FROM AGPEDIA — AGENCY THROUGH KNOWLEDGE

Effects of artificial intelligence on student learning

The effects of artificial intelligence on student learning are the measured and debated ways that AI tools, especially generative AI chatbots built on large language models (LLMs) such as ChatGPT, change what and how well students learn. Since ChatGPT's public release in November 2022, student use has grown quickly: in a fall 2024 survey, 26% of U.S. teens said they had used ChatGPT for schoolwork, double the share a year earlier.[1:1]

The evidence so far points in two directions, and the difference depends mostly on how the tool is used. In controlled experiments, general-purpose chatbots often raised students' scores while they had the tool but not what they could do without it. In one large trial they lowered later exam performance.[2:1] Tutoring systems built for teaching, which guide students instead of handing them answers, produced large gains in some trials.[3:1][4:1] Most studies are short, cover a single subject or institution, and measure outcomes within days or weeks. Long-term effects remain largely unknown.

Background

Computer-assisted instruction and intelligent tutoring systems predate generative AI by decades. General-purpose LLM chatbots differ in that students can use them for almost any task, including writing essays and solving homework problems, with no design meant to support learning. This has raised two linked concerns. The first is academic integrity: students can submit AI-generated work as their own. The second is cognitive offloading: students may hand thinking to the tool that they would otherwise do themselves, and so learn less.

Experimental evidence

The table below summarizes the main controlled studies discussed in this article.

Study Setting Participants AI condition Main result
Bastani et al. (2025) High school math, Turkey About 1,000 students GPT-4 chat vs. GPT-4 tutor with guardrails Unrestricted chat: −17% on unassisted exam; tutor: no significant harm
Kestin et al. (2025) Introductory physics, Harvard 194 students Custom GPT-4 tutor vs. in-class active learning Larger gains in less time with the AI tutor
De Simone et al. (2025) After-school English, Nigeria 759 students completed final test Teacher-guided Copilot (GPT-4) sessions +0.31 SD overall, about +0.23 SD in English
Lehmann et al. (2024) Python programming, Germany and Netherlands 176 lab participants plus 2 courses GPT-3.5 access No overall effect; exploratory evidence of harm when used to generate answers
Fan et al. (2025) English essay revision 117 university students ChatGPT-4 vs. human expert, checklist, none Better essays, no difference in knowledge gain or transfer
Kosmyna et al. (2025, preprint) Essay writing with EEG, Boston 54 adults aged 18–39 GPT-4o vs. search engine vs. no tools Weakest brain connectivity and poorest recall with the LLM

Unrestricted chatbots

The largest field experiment to date was run by Hamsa Bastani and colleagues at a high school in Turkey during the 2023–24 school year. Nearly 1,000 students in grades 9 to 11 practiced math with one of three supports: their books and notes, a ChatGPT-like GPT-4 interface ("GPT Base"), or a GPT-4 tutor given teacher-written solutions and told not to reveal answers ("GPT Tutor"). Students then took an exam without any AI. GPT Base raised practice scores by 48% but lowered exam scores by 17% compared with the control group. GPT Tutor raised practice scores by 127% and removed the harm on the exam, though it produced no measurable gain.[2:1] The authors found that students used unrestricted GPT-4 as a "crutch", often simply asking for the answer.[2:2] Students also did not perceive that copying solutions had hurt their learning.[2:3] The authors note that the study covered one subject at one school over a short period, in fall 2023, when models were less capable than later versions.[2]

Matthias Lehmann, Philipp Cornelius and Fabian Sting studied novice learners of the Python programming language. In two preregistered laboratory experiments they found no effect of LLM access on overall learning.[5:1] In exploratory analyses, students who used the LLM to generate solutions covered more topics but understood each one less well. Students who asked it for explanations understood topics better.[5:2] LLM access also widened the gap between students with low and high prior knowledge.[5:3] The study is a preprint that has not been peer reviewed, and the authors describe the substitution and prior-knowledge findings as needing replication.[5]

Yizhou Fan and colleagues randomly assigned 117 university students writing in English as a second language to revise an essay with ChatGPT-4, a human expert, a checklist tool or no support. The ChatGPT group improved their essays the most. However, the groups did not differ in intrinsic motivation, knowledge gained or ability to transfer what they learned.[6:1] The authors coined the term metacognitive laziness for the risk that learners come to rely on AI to regulate their own learning. Metacognition is the ability to monitor and direct one's own thinking.[6:2]

Purpose-built AI tutors

Greg Kestin, Kelly Miller and colleagues ran a randomized crossover trial in an introductory physics course at Harvard University in fall 2023. In a crossover trial each participant experiences both conditions. Each of 194 students took one lesson with a custom GPT-4 tutor at home and another as an in-class active learning session. The tutor was given step-by-step structure, expert-written prompts and pre-written solutions. Students learned more with the AI tutor in less time, with a median of 49 minutes, and reported feeling more engaged and motivated.[3:1] The authors estimate the effect at 0.73 to 1.3 standard deviations.[3:2] They caution that the result may not hold for topics that require combining multiple concepts or higher-order critical thinking, and that the tutor should not replace in-person teaching.[3:3]

A World Bank team led by Martín De Simone evaluated a six-week after-school English program in nine public secondary schools in Benin City, Nigeria, in mid-2024. Students typically aged 15 worked in pairs with Microsoft Copilot (built on GPT-4), guided by teachers. The program raised overall test scores by 0.31 standard deviations and English scores by about 0.23 standard deviations.[4:1] The authors estimate this is equivalent to 1.5 to 2 years of ordinary schooling.[4:2] The study is a working paper and has not been peer reviewed. Many students did not take the final test: of 1,328 students assigned, 759 completed it.[4:3] The authors report that their conclusions hold after statistical corrections for this attrition.[4]

Cognitive effects

A widely publicized 2025 preprint from the MIT Media Lab, led by Nataliya Kosmyna, recorded electroencephalography (EEG) while participants wrote essays. The 54 participants, aged 18 to 39, came from universities in the Boston area. Brain connectivity was weakest among participants using an LLM, intermediate among those using a search engine, and strongest among those with no tools.[7:1] After the first session, 15 of 18 participants in the LLM group could not correctly quote from the essay they had just written, compared with 2 of 18 in each of the other groups.[7:2] The authors describe the accumulated effect as "cognitive debt". They also note that the sample was small and geographically narrow, that only ChatGPT was tested, and that the results may not generalize to other tasks. The paper had not been peer reviewed as of its first version.[7]

Michael Gerlich surveyed 666 people in the United Kingdom and found that more frequent AI tool use was associated with lower self-reported critical thinking. The association was statistically explained by greater cognitive offloading. Younger participants, aged 17 to 25, reported the highest AI dependence and the lowest critical thinking scores.[8:1] The study is correlational, so it cannot show that AI use causes reduced critical thinking. The author acknowledges that it relies on self-reported measures and may suffer from sample bias.[8:2] A later correction replaced one table that had been duplicated in error.[8]

Academic integrity

Early fears that ChatGPT would cause a surge in cheating have not been borne out in the survey data available so far. Victor Lee, Denise Pope and colleagues compared anonymous surveys from three U.S. high schools before and after ChatGPT's release. Self-reported cheating stayed relatively stable.[9:1] Most students surveyed did not think a chatbot should be used to produce an entire paper or assignment. Many supported using one to get started on an assignment or to explain new concepts.[9:2]

Pew Research Center found similar views in its national survey of U.S. teens. 54% said using ChatGPT to research new topics is acceptable and 9% said it is not. For solving math problems, 29% said it is acceptable and 28% said it is not. For writing essays, 18% said it is acceptable and 42% said it is not.[1:2]

Policy responses

In 2023, UNESCO issued its first global guidance on generative AI in education. The guidance recommends data-privacy protections, regulatory frameworks and training for teachers. It also recommends age limits for independent use of general-purpose AI chatbots, with a minimum of 13, while acknowledging that many commentators consider 13 too young and have called for 16.[10:1]

Limitations of the evidence

The research base is young and changes quickly. Several of the most-cited studies are preprints or working papers that had not been peer reviewed, including the MIT EEG study, the Nigeria evaluation and the programming experiments.[7][4][5] Most experiments ran for days or a few weeks, involved one subject and one institution, and tested models from 2023 or 2024. A 2025 meta-analysis that reported large positive effects of ChatGPT on learning performance was retracted in April 2026 because of discrepancies in its analysis.[11:1] Its pooled estimates should not be relied on.

The open questions include:

Larger preregistered trials with longer follow-up, run across different subjects and countries, could settle these questions.

Analysis

This section contains analysis and value judgments by Agpedia contributors, based on the evidence cited above.

In the studies above, AI tended to help learning when it supported the student's own thinking and to hurt it when it replaced that thinking. Unrestricted chatbots made it easy for students to skip the effortful practice through which skills form. Across the experimental studies, students scored higher while they had the tool but gained nothing, or lost ground, once it was taken away.[2:1][6:1][5:2] Tutors that withheld answers and prompted students to reason avoided this harm and, in two trials, produced substantial gains.[3:1][4:1]

This matters for students' long-term agency. Skills and knowledge that students hold themselves give them options that depend on no particular tool. Students who rely on a chatbot to produce answers may become dependent on it without realizing it: in the Turkish trial, students did not perceive that copying solutions had hurt their learning.[2:3] Well-designed tutoring may extend agency, especially where teachers are scarce. The Nigerian program's gains were delivered in under-resourced public schools.[4:1] However, the evidence that LLM access can widen gaps between stronger and weaker students suggests that benefits will not spread evenly without deliberate design.[5:3]

  1. ^ ↗ usage-26-percent ^ ↗ acceptable-uses Sidoti, Olivia; Park, Eugenie; Gottfried, Jeffrey (2025-01-15). About a quarter of U.S. teens have used ChatGPT for schoolwork – double the share in 2023. Pew Research Center. Pew Research Center. https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/.
  2. ^a ^b ^c ↗ exam-harm ^ ↗ crutch ^a ^b ↗ no-perceived-harm ^ Bastani, Hamsa; Bastani, Osbert; Sungu, Alp; Ge, Haosen; et al. (2025-06-25). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.2422633122 https://doi.org/10.1073/pnas.2422633122.
  3. ^a ^b ^c ↗ learn-more-less-time ^ ↗ effect-size ^ ↗ limits Kestin, Greg; Miller, Kelly; Klales, Anna; Milbourne, Timothy; et al. (2025-06-03). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports. https://doi.org/10.1038/s41598-025-97652-6 https://www.nature.com/articles/s41598-025-97652-6.
  4. ^a ^b ^c ^d ↗ nigeria-gains ^ ↗ years-of-schooling ^ ↗ attrition ^a ^b De Simone, Martín; Tiberti, Federico; Barron Rodriguez, Maria; Manolio, Federico; et al. (2025-05). From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. World Bank, Washington, DC. https://documents.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf.
  5. ^ ↗ no-overall-effect ^a ^b ↗ substitute-complement ^a ^b ↗ prior-knowledge-gap ^a ^b Lehmann, Matthias; Cornelius, Philipp B.; Sting, Fabian J. (2024-08-29). AI Meets the Classroom: When Do Large Language Models Harm Learning? arXiv. https://doi.org/10.48550/arXiv.2409.09047 https://arxiv.org/abs/2409.09047.
  6. ^a ^b ↗ performance-not-knowledge ^ ↗ metacognitive-laziness Fan, Yizhou; Tang, Luzhen; Le, Huixiao; Shen, Kejie; et al. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology. https://doi.org/10.1111/bjet.13544 https://bera-journals.onlinelibrary.wiley.com/doi/10.1111/bjet.13544.
  7. ^ ↗ connectivity ^ ↗ quoting ^a ^b Kosmyna, Nataliya; Hauptmann, Eugene; Yuan, Ye Tong; Situ, Jessica; et al. (2025-06-10). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv. https://doi.org/10.48550/arXiv.2506.08872 https://arxiv.org/abs/2506.08872.
  8. ^ ↗ negative-correlation-age ^ ↗ self-report-limitation ^ Gerlich, Michael (2025-01-03). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies. https://doi.org/10.3390/soc15010006 https://www.mdpi.com/2075-4698/15/1/6.
  9. ^ ↗ cheating-stable ^ ↗ student-norms Lee, Victor R.; Pope, Denise; Miles, Sarah; Zárate, Rosalía C. (2024-12). Cheating in the age of generative AI: A high school survey study of cheating behaviors before and after the release of ChatGPT. Computers and Education: Artificial Intelligence. https://doi.org/10.1016/j.caeai.2024.100253 https://www.sciencedirect.com/science/article/pii/S2666920X24000560.
  10. ^ ↗ age-thirteen UNESCO (2023). Guidance for generative AI in education and research. UNESCO, Paris. ISBN 978-92-3-100612-8. https://www.unesco.org/en/articles/guidance-generative-ai-education-and-research.
  11. ^ ↗ retracted Wang, Jin; Fan, Wenxiang (2026-04-22). Retraction Note: The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: insights from a meta-analysis. Humanities and Social Sciences Communications. https://doi.org/10.1057/s41599-026-07310-z.
Available in