Still LearningEvidence & Practice

Better at the Task, Not Better at the Skill

Generative AI improves the work students produce. Whether it improves the student depends on which thinking it does for them. A look at the evidence.

A student hands in an analytical paragraph on Of Mice and Men. The claim is precise. The quotation is well chosen and not the obvious one. The explanation moves from language to meaning without the usual detour through plot summary. By any reasonable standard, a good paragraph.

Then you ask the ordinary follow-up question you would ask any student: why that quotation rather than the one three lines further down?

She cannot say. Not because she is evasive, but because the reasoning was never hers. She recognizes it. She cannot reconstruct it.

This is usually filed under cheating, and sometimes that is what it is. But treating it only as a conduct problem lets us avoid the harder question, which is about assessment and instruction rather than honesty. We have long used the quality of student work as a proxy for the quality of student learning. That proxy held up while producing the work required doing the thinking. It does not hold up now, and no amount of policy language about academic integrity repairs it.

The distinction we keep stepping over

Better at the task

What a student can produce while the AI is open.

Better at the skill

What the same student can do later, with the tool closed.

There are two different things we might mean when we say a tool helps students learn. The first is that students do better work while the tool is available. The second is that students are more capable afterwards, when it is not.

These are not the same measurement, and a great deal of research on AI in education measures only the first. A study can report that students using a chatbot wrote stronger essays, solved more problems and enjoyed the lesson more. All of that can be true, and none of it tells you what those students can do next Tuesday with a blank page.

This is not a criticism of the researchers. Assisted performance is cheap to measure. Independent capability is not, because measuring it means bringing students back later and taking away the thing you were studying. But it does mean the reassuring findings and the worrying ones are frequently not in conflict. They are answers to different questions, and separating them makes most of the apparent disagreement between studies disappear.

What happens when you take the AI away

What a teacher most needs to know is also the expensive thing to measure. Whether the lesson went well is visible at the time. What a student can still do once the help has been withdrawn only becomes visible weeks later, and only if someone goes back to check.

Bastani et al. (2025)Generative AI without guardrails can harm learning: Evidence from high school mathematicsPNAS, 2025 · Randomized field experiment · Grades 9–11, Turkey · ~1,000 students · Peer reviewedPractice performance rose 48% with unrestricted GPT-4 and 127% with a guarded tutor; on a later unaided exam the unrestricted group scored 17% below controls, while the tutor group was level with them.Study quality: B, strong single RCT · Confidence in the broader conclusion: moderateView original study → paid that cost. They worked with nearly a thousand students in grades 9 to 11 at a Turkish high school, randomly assigning whole classes to practice mathematics in one of three ways across four ninety-minute sessions: with an interface much like ordinary GPT-4, with a version built to push students toward attempts and explanations rather than answers, or with no AI at all.

While the AI was available it worked, and by a wide margin. Practice performance rose by 48% with ordinary GPT and by 127% with the guarded tutor.

Then the students sat an exam with no AI and nothing else to lean on.

The ones who had practiced with unrestricted GPT scored 17% below the students who had never had access to it. Not level with them, which would have been the unremarkable result, but behind a group that had spent those same weeks working the ordinary way. The guarded tutor group, despite the far larger advantage during practice, came out roughly where the control group did.

The interaction logs show what the students were actually doing. In the unrestricted condition they asked for solutions and copied them. In the tutor condition they were pushed more often into attempting, explaining and taking partial help.

The finding I keep returning to is not the headline number. The students who had learned less did not know they had learned less. Their practice had felt successful, because in the only sense available to them at the time it had been successful. That is a finding about instruction rather than conduct. We are asking adolescents to self-regulate their use of a tool that makes the experience of not-learning feel almost exactly like the experience of learning.

The unaided exam followed the practice fairly closely, so what the study shows is how a skill is acquired over a few weeks. It is often quoted as evidence of long-term cognitive decline, which is more than anyone measured.

“Students used ChatGPT” is not a description of a method

Both AI conditions in that experiment ran on the same underlying model. The model was not the intervention. The interaction design was the intervention, and it produced opposite effects on what students could do afterwards.

That should make us wary of any research question framed as AI versus no AI, and of any school policy framed the same way. The category is too coarse to predict what students will actually learn.

Build the design deliberately for learning rather than answering and the picture changes. Makransky et al. (2025)Beyond the “wow” factor: Using generative AI for increasing generative sense-makingEducational Psychology Review, 2025 · Randomized experiment · 234 Danish upper secondary students · Delayed test at 1–2 weeks · Peer reviewedA tutor designed around generative-learning principles beat ordinary ChatGPT both immediately and at the delayed test. Against conventional re-study it won immediately, but that advantage was gone by the delayed test.Study quality: B, RCT with delayed test · Confidence in the broader conclusion: moderateView original study → tried it with 234 Danish upper secondary students across four schools. After a standard lesson, students spent about half an hour either with a tutor built around generative-learning principles, with ordinary ChatGPT, or re-studying the material conventionally. Understanding was checked immediately and again one to two weeks later.

The designed tutor beat ordinary ChatGPT at both points. Against plain re-study it won immediately, and by the delayed test the advantage was gone.

I find the second half of that result more useful than the first. It is not that AI beats traditional study. It is that a carefully designed tutor can hold its own against a decent conventional alternative, while a generic chatbot performs less reliably than either. When an AI tutoring product cites research, I want to know what the comparison group was doing.

Good pedagogical intentions do not automatically become good AI design either. Fütterer et al. (2026)Enhancing school students’ self-regulated learning through generative AI support: A randomized controlled trialEducational Psychology Review, 2026 · Randomized controlled trial · 371 students, grades 7–9 · Six physics and English lessons · Peer reviewedAdding self-regulated-learning supports to students’ AI use made no measurable difference to effort, subject knowledge or strategy use when compared with ordinary ChatGPT.Study quality: B, classroom RCT, null result · Confidence in the broader conclusion: moderateView original study → added supports for self-regulated learning to students’ AI use across six physics and English lessons with 371 students in grades 7 to 9. Against ordinary ChatGPT use, the additions made no measurable difference to effort, subject knowledge or strategy use. Writing better instructions into a chatbot is not the same as changing what the student does with it.

Across the literature as a whole the picture is no steadier. Boolzen et al. (2026)Evidence of impact and interpretational limits of generative AI in STEM educationArtificial Intelligence Review, 2026 · Systematic review and meta-analysis · 85 eligible studies · Peer reviewedIndividual results scattered from substantial harm to substantial benefit, and once the tendency to publish positive findings was accounted for the pooled advantage fell close to zero. AI that substituted for the learner’s own cognitive activity was the more damaging pattern.Study quality: A, bias-adjusted meta-analysis · Confidence in the broader conclusion: moderate to highView original study → gathered 85 eligible studies of generative AI in STEM learning. Averaged in the ordinary way, the result looked encouraging. But the individual results scattered so widely that the range ran from studies showing substantial harm to studies showing substantial benefit, and once the tendency to publish positive findings was accounted for, the apparent advantage fell to something close to zero. Part of what explained the scatter was the difference between AI that augmented the learner’s own cognitive activity and AI that substituted for it, and substitution was the more damaging of the two.

An average that spans that range is not an average of one thing. It is a sign that we have been pooling interventions with almost nothing in common except the phrase used to describe them.

Using AI while learning does not automatically build dependency

The argument so far is tidier than the evidence.

Contractor and Reyes (2026)Experimental evidence on the learning impact of generative AIIZA Discussion Paper 18792 / arXiv, 2026 · Randomized experiment · Undergraduates, supervised sessions · Preprint, not yet peer reviewedStudents who had used an ordinary chatbot did better on a later unaided test, with most of the gain still present a week on. What separated them was using it to have concepts explained rather than to generate text.Study quality: C, strong design, preprint, university sample · Confidence in the broader conclusion: low to moderateView preprint → had undergraduates learn unfamiliar material and write an analytical essay in supervised sessions, with or without an ordinary chatbot, then tested them with no AI available. The students who had used it did better on the unaided test, and most of that advantage was still there a week later. Their essay quality changed little while they had access, but improved in style and relevance a week later, when they wrote unaided.

What separated those students was what they asked the tool to do. Students with AI spent less time producing sentences and more time reading and looking things up. Those who used it mainly to have concepts explained sustained their gains, while those who leaned on it to generate text lost the short-term quality advantage as soon as it was taken away.

This is a preprint rather than a peer-reviewed paper, it involved undergraduates, and it took place under supervision, so I would not offer it as evidence about a fourteen-year-old working alone at eleven at night. It points at the same distinction the school-age experiments do, approached from the other side.

None of this yet describes ordinary homework, where nobody is randomized into anything and nobody is watching. Rismanchian et al. (2026)Faster completion, less learning: Generative AI reduced study time on math problems and the knowledge they buildarXiv, 2026 · Quasi-experimental analysis · 3.2 million exercises over roughly a decade · AI exposure inferred, not observed · Preprint, not yet peer reviewedAfter ChatGPT became available, time spent on AI-susceptible problems fell about 31% among high school students and 27% among college students, with no detectable change at grade 5. Among college students the divergence vanished entirely under proctoring.Study quality: C, large-scale quasi-experimental preprint, exposure inferred · Confidence in the broader conclusion: low to moderateView preprint → examined 3.2 million mathematics exercises worked on a tutoring platform across roughly a decade, separating problems that are easy to paste into a chatbot from graph-based problems that resist it. After ChatGPT became available, time spent on the AI-susceptible problems fell by about 31% among high school students, 9% in middle school and 27% among college students, and not detectably at all in grade 5. Among the college students, that divergence vanished entirely under proctoring. On randomly assigned proctored retention items, the odds of answering correctly fell by roughly a quarter across the same period.

No short experiment can produce evidence at that scale. It is also a preprint, and the researchers never see a student use AI. They infer it from the gap between the two problem types, carefully and with falsification checks, but they infer it. So this tells us something about a population rather than proving what happened to any particular child, and the pattern across grade levels is better read as a story about who has unsupervised access than as a claim about adolescent brains.

Offloading is not the problem

“Cognitive offloading” gets used as though it were self-evidently bad. It is not. Teaching has always involved deciding what students hold in their heads and what they can reasonably put somewhere else. Notes offload memory. Calculators offload arithmetic. Dictionaries offload lexical retrieval. A well-drawn diagram offloads working memory, which is usually the entire point of drawing it.

The question was never whether to offload. It is whether the thing offloaded is peripheral to the learning goal or is the learning goal.

In English, having AI format a bibliography is almost always peripheral. If the lesson is about interpreting an ambiguous ending, having AI produce the interpretation is not peripheral, it is the lesson. Same tool, same student, entirely different consequences. In mathematics, checking arithmetic inside a modeling task is usually reasonable, since the arithmetic is not what is being taught that day; generating the solution method while that method is what the student is acquiring is a different act, though both look like “using AI for math”. In science, help organizing notes is defensible; producing the causal explanation students are meant to construct is what the lesson was for.

None of this can be settled at the level of the tool, only at the level of the task. That is inconvenient for policy writing and unavoidable for teaching.

The target thinking test

The question I now ask before designing any activity where AI might appear is short enough to be useful in a corridor conversation.

What thinking am I trying to teach here, and would the AI perform that thinking for the student?

If the answer to the second half is yes, the activity needs redesigning rather than banning, so that the thinking comes back to the student and the AI does something adjacent to it.

It helps me to sort uses into three bands. They are not rules, and the boundaries move with what I am teaching that week.

  1. AI extends the thinking

    Hints, questions, feedback after an attempt

  2. It depends on what came first

    Brainstorming, outlining, summarizing, reorganizing

  3. AI performs the target thinking

    Interpretation, procedure, argument, answer before an attempt

AI extends the thinking. It asks rather than answers. It gives a hint after an attempt, responds to work the student has already produced, quizzes, challenges reasoning, or offers a second explanation of something explained once. The cognitive work stays where it was and the student gets more of it.

It depends on what came first. Brainstorming, outlining, summarizing, reorganizing, polishing language. Whether these support or replace depends on the learning goal and on what the student did before opening the tool. Outlining is peripheral when I am teaching evidence selection and central when I am teaching essay structure. Outlining with AI after drafting your own thinking is not the same act as outlining instead of thinking.

AI performs the target thinking. It generates the interpretation, executes the procedure being taught, constructs the argument, or answers before the student has attempted anything. This is where the evidence should make us most cautious, and where the work most reliably looks best.

The uncomfortable pattern worth naming: the further down that list you go, the better the submitted work tends to be.


Before students use AI, ask five questions

1. What is the learning goal in this task? Not the topic. The cognitive operation: selecting evidence, constructing an explanation, weighing two readings against each other.

2. Which of that thinking does the student need to do themselves for the learning to happen?

3. Is the AI supporting that thinking or performing it? Hints, questions and feedback after an attempt support it. Generated answers before an attempt perform it.

4. When will the student have to do this without help? If the answer is never, I have no way of knowing whether anything was learned.

5. How will I know learning occurred, separately from how the work looks?


What this changes in practice

The most defensible immediate change is also the least technological. We can stop treating completed work as sufficient evidence of learning whenever AI could have been involved. That single move solves more problems than most AI policies do, and it requires detecting nothing.

Build friction before answers. Requiring an attempt, a prediction, or a statement of what specifically is confusing before full assistance becomes available has better school-age support than almost anything else in this area. It is not about struggle being virtuous. The generative work is the part that produces the learning, and it has to happen before something else supplies it. In practice this is a sequence rather than a rule: attempt, explain what is going wrong, ask for a hint, try again, and only then look at a full solution.

Put independent retrieval back in afterwards. Close the tool and ask the student to explain, reconstruct, solve a fresh example or teach it to someone else, and the whole activity becomes a different one. Occasionally delay that check by a week. An exit ticket and a delayed test measure different things, and the delayed one is closer to what we mean by learning.

Keep some genuinely unaided writing and problem solving, not as policing but as measurement. If I never see what a student can do without assistance, I gradually lose the ability to teach that student anything specific. That diagnostic purpose is a better argument for unaided work than academic integrity is.

Teach the distinction explicitly. Students are not being deceptive when they say they understood something. They did understand it, while it was in front of them. Nobody has given them the language to notice that this is a different state from being able to produce it.

Treat AI literacy as larger than prompting. A student who can extract an excellent answer efficiently but does not know when not to ask has learned the less useful half of the skill.

What students should understand

If I could get one idea across to a class, it would not be a rule. Using AI well is not the same as getting the best answer out of it. It is knowing which thinking you still need to do yourself, and protecting that.

The trap is specific, and it is not a character flaw. Reading a clear explanation feels almost exactly like understanding. The sentences make sense, nothing is confusing, and the fluency is real. It is just not evidence of anything, because the explanation was doing the work. The test is not whether it made sense, but what happens when it is gone.

So, after using AI on anything that matters, close it and ask:

  • Can I explain this to someone else without looking?
  • Can I reproduce the reasoning, not just the conclusion?
  • Can I do a different example?
  • If someone disagreed with this interpretation, could I defend it?
  • Can I name the part I still do not understand?

That last one is the most valuable and the one students skip. Knowing precisely where your understanding runs out is what makes the next question a good one.

What this means for parents

The evidence does not support treating AI as inherently dangerous to a child’s thinking, and there is no reason to build anxiety around occasional use.

But “did you use AI?” is not an informative question, and it tends to produce a conversation about permission rather than learning. A more useful one is: show me what you can do without it.

The patterns worth noticing are behavioral rather than technological. A child who pastes a question before attempting it. A child who cannot explain work they have submitted. A child who becomes unwilling to stay with a problem when the answer does not arrive quickly. Those map onto what the research has identified. Screen time and AI-use frequency do not.

The productive uses look different from the outside. Asking for another explanation of something half understood. Requesting a hint rather than a solution. Being quizzed. Comparing two approaches to the same problem. Practicing a language. Getting feedback on an attempt that already exists.

None of this requires supervising every interaction, which is neither realistic nor healthy. It requires making sure there are still regular occasions when a child reads, remembers, solves, explains and writes without help.

Why I would not ban it, and why I would not hand it the lesson

A blanket ban is intellectually unsatisfying and practically short-sighted. The evidence does not show that generative AI damages cognition as such. It shows that particular uses displace particular cognitive work. AI can provide explanation on demand, feedback at a scale no teacher can match, accessibility support, and a patient questioner for the student who will not raise a hand in front of thirty peers. Refusing all of that on the strength of the substitution findings is refusing the wrong thing.

But access is not surrender. AI fluency cannot substitute for knowledge, writing, reasoning, memory, interpretation, judgment and the willingness to stay with something difficult. The reason is practical as much as principled: evaluating what an AI tells you requires enough knowledge to notice when it is wrong. A student without that knowledge is not using a tool. They are trusting one.

Deciding what students should automate and what they still need to be able to do is not a technology question. It is the curriculum judgment teachers have always made, arriving with unusual force and little warning.

What we still do not know

The confident positions on both sides are running well ahead of what anyone has measured.

We do not know what happens to a thirteen-year-old who uses generative AI routinely for writing, mathematics and research across three years. Every study discussed here covers weeks at most, and the question teachers care about is developmental.

We know very little about independent writing development. The best secondary evidence on AI feedback comes from Meyer et al. (2024)Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotionsComputers and Education: Artificial Intelligence, 2024 · Randomized controlled trial · 459 upper secondary students · Peer reviewedAI-generated feedback produced real gains in the revision itself, along with better motivation and mood. No delayed, unaided writing task was included, so the study does not speak to writing development.Study quality: B for revision outcomes; the writing-development question was not tested · Confidence: low for writing developmentView original study →, who had 459 upper secondary students revise argumentative essays with or without AI-generated feedback and found real gains in the revision itself, along with better motivation and mood. It did not ask those students to write a new essay unaided weeks later, because that was not what it set out to test. So it tells us that AI feedback improves a piece of writing. It does not tell us that it improves a writer. That is not a flaw in the study but a gap in the field.

The evidence for genuine transfer is thinner still, by which I mean a different problem, in a different context, weeks later, without AI. The Danish study had to drop its harder new-question items at school level because the students found them too difficult, which is honest reporting and a fair indication of how early we are.

We do not know how any of this interacts with prior attainment. Lower-attaining students may benefit most from unlimited patient explanation, or may be the most exposed, since judging whether an AI response is any good requires the knowledge they are still building. Anyone claiming that AI closes or widens attainment gaps is ahead of the evidence in both directions.

We know almost nothing about whether students can judge their own understanding under these conditions, and that may be the most consequential gap of all. The obvious study, asking students to predict how they will do without help and then testing them without help, has barely been done.

And I teach in Norway, where there is no school-age evidence of comparable scale and design. The Danish study is the closest thing to a Nordic comparison, which is not the same as a Norwegian one. Curriculum, assessment culture, teacher autonomy and the ordinary device norms of a classroom differ enough between systems that findings should be carried across deliberately rather than assumed to travel.

The question worth keeping

The interesting question was never whether students would use AI. That was settled without our input.

The question that remains ours is what we are still asking them to learn, and whether the activity we have designed leaves that thinking with them.

Where AI removes friction from something peripheral, it is doing what a calculator or a dictionary does, and there is no reason for unease. Where it removes the retrieval, the reasoning, the interpretation or the generation that constitutes the learning, it has not saved the student time. It has done the lesson for them.

The substitution does not announce itself. It arrives as a stronger paragraph.

Research cited

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS, 122(26), e2422633122. Randomized field experiment, grades 9–11, Turkey. https://doi.org/10.1073/pnas.2422633122

Boolzen, C., Kuhn, J., Flegr, S., Rott, E.-M., Stausberg, N., & Küchemann, S. (2026). Evidence of impact and interpretational limits of generative AI in STEM education: A systematic review and meta-analysis on cognitive learning outcomes. Artificial Intelligence Review. Systematic review and meta-analysis, 85 eligible studies. https://doi.org/10.1007/s10462-026-11665-9

Contractor, Z., & Reyes, G. (2026). Experimental evidence on the learning impact of generative AI. IZA Discussion Paper No. 18792 / arXiv:2607.08849. Randomized experiment, undergraduates. Preprint, not yet peer reviewed. https://arxiv.org/abs/2607.08849

Fütterer, T., Bardach, L., Kuhn, J., Keller, S. D., & Gerjets, P. (2026). Enhancing school students’ self-regulated learning through generative AI support: A randomized controlled trial. Educational Psychology Review, 38, 42. RCT, 371 students, grades 7–9. https://doi.org/10.1007/s10648-026-10133-8

Makransky, G., Shiwalia, B. M., Herlau, T., & Blurton, S. (2025). Beyond the “wow” factor: Using generative AI for increasing generative sense-making. Educational Psychology Review, 37, 60. Two experiments; Study 2: 234 Danish upper secondary students, delayed test at 1–2 weeks. https://doi.org/10.1007/s10648-025-10039-x

Meyer, J., Jansen, T., Schiller, R., Liebenow, L. W., Steinbach, M., Horbach, A., & Fleckenstein, J. (2024). Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotions. Computers and Education: Artificial Intelligence, 6, 100199. RCT, 459 upper secondary students. https://doi.org/10.1016/j.caeai.2023.100199

Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster completion, less learning: Generative AI reduced study time on math problems and the knowledge they build. arXiv:2605.21629. Quasi-experimental analysis of 3.2 million learning interactions; AI exposure inferred, not directly observed. Preprint, not yet peer reviewed. https://arxiv.org/abs/2605.21629

Research and AI transparency

AI tools were used to help locate, organize and compare research for this article, the same kind of augmentation discussed above. They were not treated as sources. Claims about studies were checked against the original papers and reports, and the conclusions and editorial judgment remain those of Northlight Studies.