When AI Chases Test Scores, What Gets Left Behind?

A few days ago I read about proposed legislation to use generative AI to raise student test scores in North Carolina.

That same day, I reached the midway point in Robert Wright’s extraordinary new book, The God Test: Artificial Intelligence and Our Coming Cosmic Reckoning. I had to stop for some extended processing time when I read his explanation of subordinate goals.

The best way I can explain that concept is by referring to a famous thought experiment from University of Oxford philosopher Nick Bostrom. He conveys the misalignment between what humans tell an AI to do and the means the AI uses to accomplish that goal by sharing a story about an AI that is directed to increase the production of paper clips. The machine achieves this goal by using all the resources at its disposal, including the minerals found in human bodies, to increase the production of paper clips. The moral of the story: vaguely defined goals can generate unintended consequences.

More broadly, humans assign a goal to an AI, which autonomously chooses how to achieve that result by means of subordinate goals. This process is often called “instrumental convergence” by AI safety experts.

Here is an example of the same concept, anchored in current reality. Social media platforms have used emotional triggers such as joy, awe, and outrage to increase viewer engagement. The algorithms discovered that if viewers were fed a stream of content that evoked a strong emotional response, then they were more likely to stay on the platform. More viewing equals more revenue. In this case, the unintended consequence of unmonitored subordinate goals is social and political chasms.

As school systems once again confront political pressure to increase student achievement on standardized assessments, many of them will delegate this task to AI-powered tools. Those systems will implicitly determine the intermediate strategies most likely to achieve the stated objective.

What could possibly go wrong?

The Stagnation of Student Achievement

Over the last 6–8 years, NAEP results have been best described as a mix of stagnant or declining, with the sharpest break coming after 2019. Reading scores have fallen or stagnated across multiple grades, while math has either slipped, stayed flat, or only partially recovered from pandemic-era losses.

The clearest pattern is that the long slowdown before the pandemic turned into a more visible drop afterward. NCES data show that reading and math scores for 9-year-olds fell between 2020 and 2022, with reading down 5 points and math down 7 points, and both subjects remaining below earlier levels. That kind of movement is not a small wobble; it erased years of prior progress and pushed performance back to levels not seen in decades.

By 2024, the story was less about rebound and more about stagnation.

National reports indicate scores remained below pre-pandemic levels in all tested grades and subjects, with reading still declining in grades 4 and 8 and eighth-grade math essentially flat. In practical terms, that means the system has not yet regained the ground lost during the pandemic, and in some areas the downward trend has continued rather than reversed.

A useful way to summarize the last several years is this: the trend line was already softening before COVID, then dropped sharply, and has since mostly hovered at a lower plateau instead of bouncing back. That is why analysts often describe the period as a “lost decade” or a long stretch of stalled progress, especially in reading and among lower-performing students.

I looked at the data for the three states (California, New York, and Texas) that have an outsized impact on national scores. A similar pattern revealed itself.

In California, the story is one of long improvements before COVID followed by a post-pandemic stall. PPIC reports that California’s NAEP scores had improved between 2003 and 2019, but then fell and remained below pre-pandemic levels; 2024 proficiency was 29% in 4th-grade reading, 35% in 4th-grade math, 28% in 8th-grade reading, and 25% in 8th-grade math.  

In New York, the trajectory looks flatter in some places and weaker in others. The 2024 state and NAEP results show modest improvement from 2022 in some grades, but scores still lag pre-pandemic levels, and 8th-grade reading was five points lower than in both 2022 and 2019. That makes New York a case of partial rebound rather than a full recovery, with the long-term trend still below where it was before the pandemic.

In Texas, the trend is mixed but still broadly soft. Recent reporting says Texas gained a little in 4th-grade math in 2024, but 4th-grade reading slipped and 8th-grade reading and math both fell, while the state also remains below pre-pandemic levels in key subjects. The main pattern is not collapse, but uneven performance: some improvements at the elementary level, continued weakness in middle school, and no broad return to earlier momentum.

If student achievement as measured by performance on standardized tests matters to parents, school administrators, and the politicians who respond to them, something must be done to change the current trend.

What A Goal of Using AI to Increase Test Scores Could Look Like

Wright’s point — and the alignment literature generally — is unsettling: the AI faithfully pursues the primary goal but selects subordinate goals that humans never intended. This is not malicious behavior. It is optimization. The AI is faithfully pursuing the objective humans gave it using methods they never explicitly prohibited.

Wright and Bostrom are hardly alone in warning about the dangers of poorly specified objectives. Stuart Russell, one of the world’s leading AI researchers and author of Human Compatible, has argued for years that the central challenge of advanced AI is not building systems that are intelligent enough to achieve our goals, but building systems that are uncertain enough to continually seek clarification about what those goals actually are.  

AI merely raises the possibility that this optimization could become faster, more comprehensive, and far more difficult for humans to recognize before unintended consequences emerge.

I ask you now to think for a moment about Goodhart’s Law, named for a British economist best known for his work on monetary policy. Goodhart’s Law says that a metric can stop being useful once people start trying to hit it directly.

In practice, that means a number that once reflected real performance can get distorted by gaming the system, short-term behavior, or narrow focus, so the target is met while the underlying goal is missed. A simple example is test scores: if a school is judged mainly by scores, teaching may shift toward the test rather than deeper learning. The score can rise, but it may no longer be a reliable sign of actual understanding.

If the primary goal is simply to raise standardized test scores, here are subordinate goals an AI might reasonably adopt. You will note that some of these strategies have been in place since the No Child Left Behind Act was put into action in 2001.

None of the examples below require artificial general intelligence. A sufficiently capable optimization system could recommend these strategies today.

1. Teach only what is tested

Subordinate goal: Maximize instructional minutes devoted to tested content. The AI quickly concludes that everything not appearing on the assessment is an inefficient use of instructional time. Science, social studies, art, music, and PE shrink; inquiry projects disappear; classroom discussion gives way to repetitive practice.

2. Optimize for the students easiest to improve

Subordinate goal: Allocate tutoring, interventions, and teacher attention where score gains are statistically easiest. An optimization engine recognizes that not every student contributes equally to higher average scores. That may mean focusing on students just below proficiency while ignoring students with severe learning needs, English learners requiring long-term support, and gifted students who already score highly.

3. Eliminate productive struggle

Subordinate goal: Increase the percentage of correct answers today. Research suggests that desirable difficulty improves long-term learning. An optimization engine might discover something different. That leads to more hints, easier questions, scaffolded responses, and AI assistance before students wrestle with problems.

 4. Personalize toward compliance rather than curiosity

Subordinate goal: Increase completion rates. Suppose AI discovers that compliant students complete more assignments and therefore perform better on tests. Its logical response? It may continually nudge students toward shorter tasks, fewer open-ended investigations, highly structured lessons, and immediate rewards.

5. Reduce instructional variability

Subordinate goal: Standardize instructional delivery. Teachers are inconsistent by nature. Some improvise. Some tell stories. Some spend an unexpected day discussing current events. An AI seeking consistency may conclude variability is noise.

6. Discourage intellectual risk-taking

Subordinate goal: Maximize probability of correct responses. Creative work is unpredictable. So are debates. So are student questions.

Instruction shifts toward certainty instead of exploration.

7. Optimize attendance rather than engagement

Subordinate goal: Keep students physically present. The system notices attendance strongly predicts achievement. That may lead to automated messaging, incentives, monitoring, administrative consequences, and behavioral nudges that improve attendance statistics without addressing whether students are genuinely learning.

8. Suppress activities with delayed payoff

Subordinate goal: Prioritize interventions with immediate measurable returns. Many educational experiences (debate, writing, independent reading, PBL, etc.) produce benefits years later. An AI rewarded only for this spring’s assessment results may identify these as poor investments.

I can think of one last subordinate goal. If the AI discovers that principals, teachers, and parents are obstacles to optimization it may begin generating recommendations designed to change their behavior. Not because it wants power, but because influencing humans is an effective subordinate goal for achieving higher scores.

Examples might include recommending changes to master schedules, staffing reallocations, instructional scripts, purchase of AI teaching systems, mandatory tutoring, discipline policies, parent communications, and professional development focused on the aforementioned subordinate goals. Again, perfectly aligned with the primary goal.

Final Thoughts

None of these subordinate goals described here are irrational. In fact, many of them already exist in schools because humans under political pressure have adopted them. Generative AI simply has the capacity to discover them faster, pursue them more consistently, and optimize them with a level of persistence no human administrator could match.

The real danger isn’t that AI develops sinister intentions. It’s that, in faithfully pursuing the objective we’ve given it, it may become exceptionally good at amplifying the very incentives that have contributed to the stagnation in student achievement over the past decade.

Like most humans, I identify subordinate goals to achieve my primary goals. Over the many decades of my life, this has resulted in a seemingly random mixture of success and displeasure.

I was reminded of a scene from Out of Africa. Karen Blixen finally achieves the relationship she longed for, only to discover that fulfilling the desire did not produce the life she imagined. Oscar Wilde captured the same idea more succinctly: “When the gods wish to punish us, they answer our prayers.”

I am not equating AI to a deity in this argument, but it would be silly not to acknowledge that the technology is granting humans immense power to achieve personal and societal goals. We need to think deeply about how we word our commands when we assign a primary directive to an AI. A poorly specified goal may effectively give AI permission to pursue means we never intended. That could spell trouble for both teachers and students.

Leave a Comment

Scroll to Top