Skip to main content
Evaluator last updated June 22, 2026.

Overview

The Math Alignment evaluator checks a math question against a standard’s individual learning components — not just against the standard’s label. It reports which of a standard’s components the question actually measures.
The Math Alignment evaluator judges whether a question is the right math for a standard. The Math Visual Correctness evaluator (coming soon), judges whether a math visual is mathematically correct.

At a glance

The evaluator was built and validated using the model and temperature below (other configurations will produce different results and may have lower accuracy):

Getting started

Follow the Quickstart to start using this evaluator:

Inputs

Inputs must be de-identified. Do not submit student PII or any regulated or sensitive personal information.
The evaluator supports 3 modes:
Use Batch and By-grade evaluation to surface which standards (and which learning components) are fully, partially, or not covered by a given question bank.

Output

The evaluator reduces alignment to a binary judgment (plus rationale) per learning component, and is not validated for grading, assessment, or placement decisions.Treat outputs as directional signals, and keep a human in the loop – especially for borderline cases.

Interpreting results

The evaluator assesses standards alignment based on how many of a standard’s learning components a question meets (i.e., Aligned count / Total count). Users should interpret the counts while keepign in mind the learning components a question was meant to target. An Aligned count of 2 out of a Total count of 5 is ambiguous information on its own. If those 2 Aligned count learning components include the ones a question was meant to target, you can conclude the question is aligned. Otherwise, you can conclude the question is not aligned.
Example: A question that asks students to find the area of a rectangle by multiplying its sides’ lengths is commonly tagged to Common Core 3.MD.C.7. However, that question meets only one of the 4 learing components that make up that standard.At the parent-code level the question looks aligned; at the learning-component level, it covers a quarter of the standard.

Accuracy and validation

This evaluator is provided as Early access. Comprehensive accuracy measures are still evolving, and validation testing is ongoing.
The prompt was optimized using GEPA ↗ via DSPy ↗ on a stratified split of 2,011 question-learning component pairs (724 train / 482 validation / 805 test). These were drawn from Illustrative Mathematics v.360 ↗ cool-down questions (CC-BY-NC 4.0) and annotated by 3 human experts. The GEPA-optimized prompt was evaluated on the held-out test split, and separately reviewed by math experts on a sniff test of brand-new questions.
The evaluator was designed and validated on Illustrative Mathematics 360 cool-down questions, which are not evenly distributed across grades K–12 — performance may vary by grade. The evaluation set also may not fully represent the range of math inputs developers could submit, so brand-new or unusual inputs carry more risk of poor results.

Evaluator release history