Thousands of STAAR reading results improved after humans rescored students’ written answers. Most of the initial grading was done by an automated engine.
The number of requests to rescore students’ STAAR tests nearly tripled from last year as more Texas schools sought reviews of written responses on the reading exams, thousands of which were initially graded by computers.
The Texas Education Agency began using an automated scoring engine in December 2023 to grade open-ended questions on the State of Texas Assessments of Academic Readiness, including on the reading and English tests, with humans manually reviewing a quarter of exam scores.
In spring 2024, the state had 2,758 requests to regrade open-ended response scores on reading exams. Then, 21,620 requests last year. This spring, that jumped to 64,489.
Meanwhile, the number of corrections doubled compared to the previous year. The state improved more than 13,300 reading scores after the reviews. Of those, humans initially assigned grades to 42% while computers graded 58%. Humans conducted all re-scores.
Allison Matney, director of the Texas Center for School Accountability, said historically, it was rare for district administrators to request STAAR rescoring. She attributed the increase to greater awareness about the automated scoring and review process.
“That traction of superintendents talking to each other — (and) media continuing to talk about it — made more districts aware that not only can they do this process, but actually they should, and they need to,” Matney said.
The state’s use of automated scoring for open-ended responses was initially criticized after a spike in students earning zeros on writing prompts in the first round of exams graded by computers. Agency officials, however, pointed to other causes for the low scores in 2023, including that many students were re-taking the exam.
State officials have stressed that automated computer scoring is different from generative artificial intelligence, such as ChatGPT or Google’s Gemini, as the program does not “learn” from one response to the next.
Agency spokesperson Jake Kobersky said scoring essays by hand would require four to five times the number of human scorers and cost an additional $15 million to $20 million per year. It would also result in delayed results for districts, he said.
“TEA uses a hybrid model where human experts drive the scoring. … The result is a system that keeps human judgment at the center of scoring, applies it where it matters most, and delivers results to districts and families faster than hand-scoring every response would allow,” Kobersky said.
More than 3.2 million Texas students took the reading STAAR this spring, which includes third through eighth graders, as well as those who took the high school level English I and English II end-of-course assessments. In total, less than .5% of total open-ended responses on reading exams saw changes.
TEA officials began incorporating automated scoring after the state redesigned the STAAR to include fewer multiple-choice questions and more open-ended questions — which are also known as constructed responses.
To develop the computer scoring, the TEA gathered thousands of responses that were previously scored by people. From the sample, the computer learned the characteristics of responses, and it was programmed to assign the same scores as a human.
More than 250 public charter networks and school districts, including Cypress-Fairbanks, Dallas, Katy and Northside, submitted rescoring requests this year. Nearly 90% had at least one submitted reading exam see a score increase, according to agency data.
Source: Texas Tribune BY Megan Menchaca
Photo: A Nimitz Middle School student raises their hand during class on Sept. 13, 2023 in Odessa. Eli Hartman/The Texas Tribune





