Skip to main content
Evaluator last updated June 24, 2026.

Overview

The Appropriateness evaluator assesses whether a piece of feedback correctly identifies whether a student’s response has met the task goal and thus does not require further revision. The evaluator considers:
  • Accurate task assessment (whether the response meets the task’s relevant requirements)
  • Revision-need identification (distinguishing between feedback that should direct revision and feedback that should signal “move on”)
  • Appropriate signal when the task is complete (signaling the student can move on when the goal is met)

At a glance

The evaluator was built and validated using the model and temperature below (other configurations will produce different results and may have lower accuracy):

Getting started

Follow the Quickstart to start using this evaluator:

Inputs

Inputs must be de-identified. Do not submit student PII or any regulated or sensitive personal information.
Example input

Output

Example output

Interpreting results

Accuracy and validation

This evaluator is provided as Early access. Reported metrics come from a small held-out test split (20 examples) with wide confidence intervals and should be read as directional. Validation testing is ongoing.
We assessed performance against Quill.org ↗ classroom writing data (52 labeled pairs; 16 train / 16 validation / 20 test) — expert-annotated student-response and teacher-feedback pairs labeled by Leanlab Education ↗ using the Productive Coaching rubric.
On this dimension, GEPA optimization improved GPT-5.4 substantially over its naive baseline on the held-out test set (+20 points accuracy, +25 points macro-F1) — the largest optimization gain in the suite.

Evaluator release history