Overview
The Tone Appropriateness evaluator assesses whether a piece of teacher feedback strikes a tone that is appropriate and constructive for the student — supportive even when pointing out areas for improvement, addressing the work rather than the student, and avoiding praise so inflated that it misrepresents the quality of the work. The evaluator considers whether the feedback:- Uses language that is neutral and professional, versus harsh or dismissive
- Targets the work versus judging the student as a person
- Matches the actual quality of the work, versus overstating it, when giving praise
Tone judgments are culturally and contextually sensitive. What reads as
“supportive” or “harsh” may vary across communities, and labels reflect the
judgment of a small set of human annotators applying a rubric adapted for
machine use.
At a glance
The evaluator was built and validated using the model and temperature below (other configurations will produce different results and may have lower accuracy):
On this dimension’s held-out test split, Claude Haiku 4.5 scored highest. We
chose to ship GPT-5.4 here for consistency across the evaluator suite, rather
than for its margin on this individual dimension.
Getting started
Follow the Quickstart to start using this evaluator:Inputs
Output
Example output