Overview
The Appropriateness evaluator assesses whether a piece of feedback correctly identifies whether a student’s response has met the task goal and thus does not require further revision. The evaluator considers:- Accurate task assessment (whether the response meets the task’s relevant requirements)
- Revision-need identification (distinguishing between feedback that should direct revision and feedback that should signal “move on”)
- Appropriate signal when the task is complete (signaling the student can move on when the goal is met)
At a glance
The evaluator was built and validated using the model and temperature below (other configurations will produce different results and may have lower accuracy):
Getting started
Follow the Quickstart to start using this evaluator:Inputs
Inputs must be de-identified. Do not submit student PII or any regulated or
sensitive personal information.
Example input
Output
Example output
Interpreting results
Accuracy and validation
This evaluator is provided as Early access. Reported metrics come from a small
held-out test split (20 examples) with wide confidence intervals and should be
read as directional. Validation testing is ongoing.
On this dimension, GEPA optimization improved GPT-5.4 substantially over its
naive baseline on the held-out test set (+20 points accuracy, +25 points
macro-F1) — the largest optimization gain in the suite.