Skip to main content

The problem

Constructs like critical thinking, argumentation, or reasoning are core to what schools want students to develop, and they shape rubrics, feedback, and instruction. But rating student work against a construct like this is harder than it looks:
  • Construct is rarely a single skill
    • Most constructs bundle several indicators (e.g., synthesizing sources, addressing counterarguments, drawing conclusions) that don’t always move together
    • Essay can be strong on one indicator and weak on another – a single holistic score hides that
  • Construct isn’t the same as writing quality
    • Fluent, well-organized writing can be thin on reasoning, while rough writing can also carry real reasoning
    • Rating must isolate the construct itself, not assess general essay quality
  • Even human experts don’t always agree
    • Reliable rating takes calibration: shared rubrics, normalizing sessions, and consensus across raters
    • Some indicators remain difficult to rate consistently, even after calibration – that ceiling must be reported, not hidden
Asking human experts to rate every student’s work against every indicator, across a full class or cohort, does not scale.

What we’re building

Our Durable Skills evaluators rate student work against a construct’s indicators, using a rubric developed and calibrated with subject-matter experts. Our evaluators judge whether student work demonstrates a construct at the indicator level:

Quickstart

Run an evaluator in the Evaluators playground, a Python notebook, or with the SDK.

Critical Thinking evaluator

Rate a student essay for critical thinking, with evidence quotes and rationale for every indicator.