The problem
Constructs like critical thinking, argumentation, or reasoning are core to what schools want students to develop, and they shape rubrics, feedback, and instruction. But rating student work against a construct like this is harder than it looks:- Construct is rarely a single skill
- Most constructs bundle several indicators (e.g., synthesizing sources, addressing counterarguments, drawing conclusions) that don’t always move together
- Essay can be strong on one indicator and weak on another – a single holistic score hides that
- Construct isn’t the same as writing quality
- Fluent, well-organized writing can be thin on reasoning, while rough writing can also carry real reasoning
- Rating must isolate the construct itself, not assess general essay quality
- Even human experts don’t always agree
- Reliable rating takes calibration: shared rubrics, normalizing sessions, and consensus across raters
- Some indicators remain difficult to rate consistently, even after calibration – that ceiling must be reported, not hidden
What we’re building
Our Durable Skills evaluators rate student work against a construct’s indicators, using a rubric developed and calibrated with subject-matter experts.
Our evaluators judge whether student work demonstrates a construct at the indicator level:
Related topics
Quickstart
Run an evaluator in the Evaluators playground, a Python notebook, or with
the SDK.
Critical Thinking evaluator
Rate a student essay for critical thinking, with evidence quotes and
rationale for every indicator.