Skip to main content

The problem

Research ↗ shows that students who consistently engage with complex texts are more likely to succeed in college and beyond. Yet despite their importance, complex texts often remain absent from classrooms.
  • Quantitative measures of text complexity (e.g., Lexile or Flesch-Kincaid) are useful, but limited
  • Qualitative measures are more accurate, but also more labor-intensive to assess
As AI-generated texts enter the classroom, educators risk using content that looks grade-appropriate on the surface, but fails to meet the deeper demands of literacy development.

What we’re building

Instead of giving a single complexity score, our literacy evaluators assess text across multiple qualitative dimensions. They are anchored in Student Achievement Partners (SAP)‘s Qualitative Text Complexity Rubric for Informational Text ↗, giving you:
  • Fine-grained data to ensure quality generated texts
  • Actionable insights into why a text may be complex or not complex enough and how to best scaffold it for students
Explore our Literacy dataset for the benchmark data that our literacy evaluators use to assess text complexity.

Scope and limitations

Literacy evaluator outputs should not be used for high-stakes applications like grading, assessment, or placement decisions without human review.
Remember that LLM scores can vary across runs, especially on borderline cases. We recommend keeping a human in the loop and treating outputs as directional signals vs. definitive judgments.

Literacy dataset

Explore the expert-annotated benchmark behind literacy evaluators.

SDK API Reference

Integrate literacy evaluators into your TypeScript or Python project.

Feedback evaluators

Explore evaluators that assess the quality of coaching feedback.

Standards evaluators

Explore evaluators that assess content alignment to academic standards.