Skip to main content
Evaluator last updated August 27, 2026.

Overview

The Grade Level Appropriateness evaluator assesses whether AI-generated text is suitable for independent reading at a specified grade band. The evaluator considers:
  • Flesch-Kincaid grade level
  • Word count
  • Text structure – Organization complexity, connections between ideas, text features
  • Language features – Vocabulary, sentence complexity, figurative vs. abstract
  • Purpose – Explicitly vs. not explicitly stated, concrete vs. abstract
  • Knowledge demands – Discipline-specific knowledge, references, allusions
  • Student background knowledge – What students at a given grade level would already know

At a glance

The evaluator was built and validated using the model and temperature below (other configurations will produce different results and may have lower accuracy):

Getting started

Follow the Quickstart to start using this evaluator:

Inputs

Output

Interpreting results

Evaluator release history

Literacy evaluators

Explore other evaluators that assess qualitative text complexity.

Literacy dataset

Explore the expert-annotated benchmark behind literacy evaluators.

Quickstart

Run an evaluator in the Evaluators playground, a Python notebook, or with the SDK.