Skip to main content
← Back to Blog
·14 min read·By Priscilla Han

The Homework Trap: When Better Marks Mean Less Learning

A CEPR study of 26,811 Chinese students found that using AI pushed homework scores up about 18% while exam scores fell up to 24%. The cause is homework outsourcing. Here is what it means, what learning science says about why, and a simple test to catch it at home before an exam does.

AICritical ThinkingLearning ScienceData

The Short Answer

A strong homework grade has always been reassuring news. It meant the work was getting done and, usually, that the learning was happening underneath. A new working paper from a team of economists says that reassurance is no longer safe. When students used AI to do their homework, their homework scores rose by about 18 percent, while their exam scores fell by roughly a fifth within six months, and by as much as a quarter on the high-stakes entrance exams. The students whose homework looked best were often the ones learning the least. Below is why that happens, what four decades of learning research predicted it would, and a short test you can run at home this week to catch it early.

+18% homework, −20% exams

What outsourcing homework to AI did to Chinese secondary students: homework scores up about 18%, monthly exam scores down about 20% within six months

Strömberg, Lei & Wu (2026), CEPR Discussion Paper 21577

The study, and the line that goes the wrong way

What the researchers measured

In June 2026, economists David Strömberg, Victor Lei and Yanhui Wu released a working paper through the Centre for Economic Policy Research, "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education" (Discussion Paper 21577). They followed 26,811 students in grades 7 to 12, roughly ages 12 to 18, across 30 months of panel data. That scale and duration matter. This is not a lab exercise with fifty undergraduates over an afternoon. It is tens of thousands of real students, doing real coursework, tracked over two and a half years as generative AI arrived in their lives. The Economist highlighted the result in a Graphic Detail piece, "Does AI stop children from learning?", on 18 August 2026.

The reversal

The headline is not the blunt "AI is bad for children." It is subtler, and far more useful to a parent. When students used AI for their homework, the paper reports, three things moved together: homework scores rose about 18 percent, the time spent on that homework fell about 30 percent, and monthly exam scores dropped around 20 percent within six months. On the entrance exams that actually gate a child's future, the fall was steeper still, between 18 and 24 percent.

What using AI didChange for AI users
Homework scoresabout +18%
Homework completion timeabout −30%
Monthly exam scores (within six months)about −20%
High-stakes entrance-exam scoresabout −18% to −24%

Source: Strömberg, Lei and Wu (2026), CEPR Discussion Paper 21577. It is a working paper, not yet peer-reviewed.

The Economist's chart makes the reversal vivid. It plots exam score against homework score, each indexed so that the average student who never uses AI sits at 100. For students who never used AI, the two rise together: better homework, better exams, the relationship we all grew up with, reaching about 108 at the top of their range. For AI users the line starts the same way, peaks near a homework score of 105, and then bends downward. Among the very highest homework scorers who used AI, average exam results fall to somewhere around 62. Those figures are approximate, read from the published chart, but the shape is the whole point.

For AI users, exam scores rise with homework only to a point, then fall

Average exam score by homework band, students using AI (no-AI average = 100). Students who never used AI kept rising instead, to about 108 at the top of their range. Approximate, read from the Economist figure.

Homework ~100exam ~96
Homework ~105 (peak)exam ~97
Homework ~110exam ~95
Homework ~120exam ~78
Homework ~130exam ~62

Strömberg, Lei & Wu (2026), CEPR DP21577

It was outsourcing, not simply "using AI"

The most important number in the paper is not any of the ones above. It is this: about 80 percent of the students using AI were, in the authors' term, outsourcing their homework, finishing far faster than before while their marks climbed. Those were the students who lost the most at exam time. The AI users who kept a normal working pace saw only small losses. So the damage was not spread evenly across everyone who touched the technology. It concentrated in exactly the children whose homework looked most impressive. Hold onto that, because it is what makes the finding actionable rather than merely alarming.

Why homework stopped being a signal

Homework was a proxy, and the proxy broke

Homework was never valuable in itself. No school offers a place because a child's Tuesday-night worksheets were tidy. Homework mattered because it was a reliable proxy: a cheap, everyday signal standing in for something we care about but cannot see directly, which is whether the learning is happening. If the homework was good, the learning was almost certainly there, because the only way to produce good homework was to do the thinking.

Economists have a name for what goes wrong when a proxy is leaned on too hard. It is Goodhart's Law, after Charles Goodhart's 1975 observation that a statistical regularity tends to collapse once it is used for control. The familiar phrasing, "when a measure becomes a target, it ceases to be a good measure," is Marilyn Strathern's later 1997 generalisation of the idea. Generative AI turns homework into a target a student can hit without doing the learning the measure was invented to capture. That is why the line does not simply flatten, it inverts. The students who could offload the most did, so a very high homework mark increasingly flagged the heaviest reliance on the tool rather than the deepest understanding.

The fingerprint is speed

The completion-time finding is what turns this from a plausible story into a measurable one. Work that used to take an evening now took minutes, and the marks went up anyway. Speed, not the marks, is the tell. A child who is genuinely learning with a tool usually spends about the same time, or more, because they are checking, retrying, and questioning. A child who is outsourcing gets faster and higher at the same moment. That combination, quicker and better, is the pattern the researchers used to separate the students who crashed from the ones who did not, and it is a pattern you can watch for at your own kitchen table.

The learning that AI can quietly remove

None of this is new to the people who study how memory works. The study confirms, at enormous scale, what four decades of cognitive research would have predicted. Four findings are worth knowing, because together they explain the whole curve.

Effort is the mechanism, not the obstacle

The struggle to retrieve a fact, to reconstruct a method you half-remember, to sit with a problem before the answer arrives, is not friction getting in the way of learning. It substantially is the learning. Robert and Elizabeth Bjork call the conditions that create this useful struggle "desirable difficulties" (Bjork and Bjork, 2011): things like spacing practice out, mixing problem types, and testing yourself, all of which make learning feel slower and harder in the moment while making it far more durable. A tool that removes the difficulty removes the mechanism.

Being tested is how knowledge sticks

Related, and one of the most robust results in the field: the act of retrieving something from your own memory, rather than rereading it, is what makes it stick. In a classic experiment, Roediger and Karpicke (2006) showed that students who tested themselves on material remembered dramatically more a week later than students who simply restudied it, even though the restudiers felt more confident (the testing effect). Homework, done properly, is retrieval practice. Homework handed to a machine is not retrieval of anything.

Offloading to a tool means remembering less

There is also direct evidence about what happens when we expect a tool to hold information for us. Sparrow, Liu and Wegner (2011) found that when people believed facts would remain available on a computer, they remembered the facts themselves less well, while remembering where to find them ("Google effects on memory", often called cognitive offloading). AI is that effect with the volume turned up. It does not just store the fact, it performs the entire task, so there is even less reason for the mind to hold onto anything.

The child feels fluent, and the feeling misleads

The cruelest part is that outsourced learning feels like learning. The child reads a clean AI answer, follows it, and feels a warm sense of "yes, I understand this." Psychologists call this the illusion of fluency, and it has real consequences: Dunlosky and Rawson (2012) found that students who overestimate what they have learned go on to study less and achieve less. Overconfidence is not harmless. It actively suppresses the effort that would have closed the gap. A tool that manufactures fluency without understanding is, from this angle, manufacturing overconfidence, which is worse than manufacturing nothing.

It is not only children

If this were purely about developing brains, we might hope children grow out of it. The early evidence says otherwise. A 2025 study of knowledge workers found that higher confidence in generative AI was associated with less critical thinking, and that the tool tended to shift effort toward checking its output rather than doing the original thinking (Lee and colleagues, 2025). Different population, same shape. The habit of letting the tool think is not a phase. It is a skill that atrophies with disuse, at any age.

The real divide is not AI versus no AI

This is the nuance the study hands parents, and it is the most important line in this article. The children who suffered were not simply the ones who used AI. They were the ones who used it to replace their own thinking.

At one end is AI as a tutor. The student gets stuck, asks it to explain a concept a different way, tries the next step, checks their own reasoning, asks why a wrong answer was wrong. The effort stays with the child; the tool removes confusion, not thinking. Used this way, AI can be a patient, always-available teacher of a kind most families could never otherwise afford, and there is real promise in that.

At the other end is AI as a ghostwriter. Paste the question, copy the answer, move on. The homework is done and the learning never started. This is the behaviour the completion-time data caught, and it is the behaviour that showed up months later in the exam hall.

One honest caveat. This is a working paper, not yet peer-reviewed, and it describes what happened in one country's schools, not a law of nature. But it does more than note a correlation. By identifying the outsourcing pattern, fast completion times alongside inflated homework marks, it separates the students who lost the most from those who did not, and the separation lines up precisely with what the learning science above would predict.

What to do this week: the Tutor Test

You do not need to ban AI, and for most families that would not work anyway. You need a way to see through the homework grade to the learning underneath, the way the exam eventually will, but sooner and more gently. Here is the framework I share with the families I work with. Four checks, none of which requires you to understand the subject yourself.

1. The explain-it-back test. After a piece of homework is done, ask your child to explain one answer out loud, with the screen and notes closed, in their own words, including why it is right. A child who did the thinking can usually manage this, even clumsily: "I divided both sides because I needed x on its own, and I checked it by putting the number back in." A child who outsourced it produces a confident opening sentence that dissolves the moment you ask a "why" or "what if": "It is just the formula" with nothing underneath. This is the closest thing you have to a home exam, and it is really just retrieval practice by another name.

2. Watch the gap, and the clock. The study's whole insight is a gap between strong homework and weaker exams, and its tell is speed. Line up recent homework marks against quiz and test marks: if homework is consistently strong but anything closed-book and timed lags, that divergence is your early warning. And notice how long the work takes. Homework that suddenly gets done in a fraction of the usual time, at the same or higher marks, is the single clearest sign of outsourcing, because it is the exact fingerprint the researchers used.

3. Ask to see the mess. Learning is messy: crossings-out, wrong turns, a first attempt that failed. Outsourced work arrives suspiciously clean, a finished product with no visible struggle. Ask, without accusation, to see the working and the drafts. Their absence is informative.

4. Change the question, not the tool. Instead of "did you use AI," which invites a defensive yes or no, ask "how did it help." The answer sorts tutor from ghostwriter instantly. "It explained why my method was wrong so I could redo it" is a tutor. "It gave me the answer" is a ghostwriter. You are not policing the tool. You are teaching your child the difference between getting unstuck and getting bailed out, a distinction they will need for the rest of their working life.

The goal of all four is the same: to make the invisible visible. The exam does this too, but once, late, and with high stakes. You can do it every week, cheaply, kindly, and in time to change course.

What it means for the decisions you are making

If you are choosing a school, this study reframes a question worth asking on every tour. It is one we now raise with families at BrightKey when we help them weigh a shortlist. Not "do you use technology in the classroom," which every school answers yes. Ask instead how the school tells the difference between a child who is learning with AI and one who is outsourcing to it, how it assesses understanding in ways a tool cannot fake, and whether closed-book, in-person work still carries real weight. A school that has thought hard about that is protecting your child's learning. A school that simply markets itself as "AI-powered" may, without meaning to, be celebrating faster homework and thinner learning.

This is the same tension we explored from a different angle in Is AI Making Our Kids Smarter, Or Just Faster?. The tools change; the underlying question does not. Output is easy to inflate. Understanding is not, and it is the only thing that travels with your child into the exam hall, the interview, and the first job.

We are parents too, and we know how tempting it is to read a strong homework grade as a job well done and move on. The most useful thing this research does is take that comfort away, and hand us something better in its place: a reason to look one level deeper, and a simple way to do it.

If you would like a second pair of eyes on your own situation, whether a school shortlist, a subject that is not adding up, or how to set healthy AI habits at home, you are welcome to message me directly on WhatsApp at +65 9656 2770.

Sources and caveats. The homework, completion-time and exam figures are the paper's reported results: David Strömberg, Victor Lei and Yanhui Wu, "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education", CEPR Discussion Paper 21577, June 2026, a working paper not yet peer-reviewed. The exam-versus-homework chart values are approximate, read from the figure published by The Economist. Learning-science references: Roediger and Karpicke (2006) on the testing effect; Bjork and Bjork (2011) on desirable difficulties; Sparrow, Liu and Wegner (2011) on cognitive offloading; Dunlosky and Rawson (2012) on overconfidence; Goodhart (1975) and Strathern (1997) on the law of measures becoming targets; Lee and colleagues (2025) on AI and critical thinking in adults. Treat the central finding as a strong, well-identified signal worth acting on, not the final word.

Need guidance on this topic?

Book a free 30-minute consultation with Priscilla.

Get in Touch