About the dataset
The Literacy dataset provides text-complexity annotations for the CLEAR (CommonLit Ease of Readability) Corpus by literacy experts and qualified educators. It is the benchmark data that Evaluators use to assess literacy levels in AI-generated text. We are sharing it as a new resource for the learning science community to help address the need for more high-quality text-complexity datasets and to complement existing work in this area. The CLEAR Corpus was produced by CommonLit in collaboration with Georgia State University ↗ and released in December 2021. It comprises nearly 5000 publicly available excerpts, each mapped against dimensions including Flesch-Kincaid and BT Easiness (Bradley-Terry coefficient based on teacher ratings of the text). We expanded the dataset by scoring a subset of rows for text complexity dimensions found in Student Achievement Partners’ Qualitative Text Complexity rubric (SAP) ↗. Our initial release in September 2025 focuses on Grades 3 and 4 across sentence structure and vocabulary, but we plan to expand to all grades and all dimensions of text complexity assessed through SAP’s Qualitative Text Complexity rubric. Thanks to Student Achievement Partners and to Achievement Network for their contributions in helping us assemble this annotated data.Our process
Our process for producing annotated data is as follows:- Filter the CLEAR corpus to an approximate grade 3-4 range using Flesch Kincaid Grade Level.
- Partner with literacy experts from SAP (Student Achievement Partners) and ANet (Achievement Network) to score against text complexity dimensions on SAP’s Qualitative Text Complexity rubric.
- With SAP and ANet, establish a gold set of ~80 examples per grade with representation across the four tiers of text complexity: slightly, moderately, very, and exceedingly complex.
- Use the gold set to test and qualify a cohort of educators with at least 2 years of experience teaching ELA at the corresponding grade level.
- Produce a minimum of 200 rows (at 50 per complexity tier), calibrating annotator scores using the Dawid-Skene method.
Columns
Last updated September 23, 2025This list will be updated as we incorporate more text complexity dimensions.
Receiving Not Scored for a given column means that the text was not
annotated for that column.
Limitations
- Annotator coverage per item is limited; reported precision and agreement reflect this coverage.
- Annotations in future updates will improve statistical reliability by increasing confidence, reducing variance, and stabilizing borderline cases.
- Flesch–Kincaid grade estimates (based on sentence and word length) are heuristic and do not capture qualitative factors (e.g., conceptual difficulty, vocabulary sophistication, thematic maturity).
- They are recorded as metadata and may be referenced in the evaluator prompt as one of multiple signals.
- They are not the sole determinant of evaluator labels.