
The short answer
Students who could use an unrestricted AI chatbot while practising maths problems did noticeably better on that practice, then did worse on a later exam taken without any AI help, compared with students who had practised without a chatbot at all. A version of the same tool that offered hints rather than answers avoided most of that later harm.
What the evidence says
Researchers Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakçı and Rei Mariman ran a field experiment with close to a thousand high school maths students, randomly assigning classes to practise problems with different levels of AI access: a standard chatbot that could simply provide answers, a version built to give teacher-style hints rather than solutions, or no chatbot at all. The working paper, first posted on SSRN on 18 July 2024, and its later peer-reviewed version in the Proceedings of the National Academy of Sciences, published 25 June 2025, report that students using the unrestricted chatbot showed large gains during practice sessions but performed worse on a subsequent test taken without AI assistance than students who had practised unaided. Students using the hint-based, teacher-designed version of the tool did not show that same later decline to the same degree, suggesting the design of the AI's response, not the presence of AI itself, drove the harm.
For context
The mechanism the authors describe is a familiar one in learning research under a new name: a tool that hands over the answer can boost performance while it is available and leave less durable understanding behind once it is removed, because the student practised retrieving an answer rather than working through the reasoning. This was a field experiment with random assignment to conditions, which supports a reasonably confident causal reading of the comparison between groups, though it covers a specific age group, subject and time-limited study period rather than every possible way a student might use a chatbot. PNAS issued a correction to the published paper in August 2025, a reminder that even a peer-reviewed, widely discussed study can need a subsequent formal amendment, and that the underlying comparison between unrestricted and hint-based AI access is what has drawn the most attention from other researchers since.
A practical next step
Before allowing open chatbot access for homework practice, you could ask whether the tool is configured to prompt and hint, the way a teacher would, rather than simply supplying finished answers.
- Does this tool give the answer immediately, or does it ask a guiding question first?
- Is the child's performance being checked without AI access at all, at some later point?
- Would a teacher-supervised, hint-only setting suit this subject better than open access?
The distinction this study draws, between a chatbot that teaches and one that simply finishes the work, is likely to matter more for learning outcomes than the blunter question of whether to allow AI tools in the first place.
Sources & reading trail
SSRN working paper record confirms the original 18 July 2024 posting of the field experiment and its title.
Source published: 18 July 2024 · Retrieved: 16 September 2026
Peer-reviewed PNAS abstract states the field experiment's nearly 1,000-student sample, the practice-versus-exam performance pattern, and that safeguarded hints reduced the harm.
Source published: 25 June 2025 · Retrieved: 16 September 2026
Studies and official documents establish the record; the short answer and the next step are Screens & Childhood editorial interpretation. This retrospective draft does not imply the site published on the event date.