About the dataset
The Literacy dataset contains high-quality text complexity annotations for the CommonLit Ease of Readability (CLEAR) Corpus by literacy and education experts. The Corpus was produced by CommonLit in collaboration with Georgia State University ↗ and released in December 2021. It comprises nearly 5000 publicly available excerpts, each mapped against dimensions including Flesch-Kincaid and Easiness (Bradley-Terry coefficient based on teacher ratings of the text). We expanded the dataset by scoring a subset of rows for dimensions in Student Achievement Partners ()‘s Qualitative Text Complexity Rubric for Informational Text ↗. Our Literacy evaluators use this dataset as a benchmark when assessing AI-generated content. We hope edtech developers can use it to complement their own literacy work as well.Our initial release in September 2025 focuses on Grades 3 and 4 across
sentence structure and vocabulary, but we plan to expand to all grades and all
dimensions of text complexity assessed through
’s Qualitative Text
Complexity rubric.
Our process
Our process for producing annotated data is as follows:- Filter the corpus to an approximate grade 3-4 range using Flesch Kincaid Grade Level.
- Partner with literacy experts from and Achievement Network () to score against dimensions in SAP’s Rubric.
- With SAP and ANet, establish a gold set of ~80 examples per grade with representation across the 4 tiers of text complexity (Slightly, Moderately, Very, and Exceedingly complex)
- Use the gold set to test and qualify a cohort of educators with 2+ years of experience teaching English Language Arts () at the corresponding grade level.
- Produce 200+ rows (50+ per complexity tier), calibrating annotator scores using the Dawid-Skene method.
Columns
Last updated September 23, 2025This list will be updated as we incorporate more text complexity dimensions.
Receiving Not Scored for a given column means that the text was not
annotated for that column.
Limitations
- Annotator coverage per item is limited; reported precision and agreement reflect this coverage.
- Annotations in future updates will improve statistical reliability by increasing confidence, reducing variance, and stabilizing borderline cases.
- Flesch–Kincaid grade estimates (based on sentence and word length) are heuristic and do not capture qualitative factors (e.g., conceptual difficulty, vocabulary sophistication, thematic maturity).
- They are recorded as metadata and may be referenced in the evaluator prompt as one of multiple signals.
- They are not the sole determinant of evaluator labels.