Abstract Since generative AI entered educational institutions, discussions within the education sector about its impact have largely centered on the question, 'Does AI harm student learning?' However, the actual situation requires more careful examination. Two rigorous studies published between 2025 and 2026 superficially appear to support the conclusion that AI causes harm. Yet, upon closer inspection of the variations within groups, the true insight does not lie in the average difference between experimental and control groups, but in why such a wide range of academic outcomes exists among students who all used AI. This observation suggests that the question itself may be misplaced: it assumes AI is a singular (all-or-nothing) independent variable, while also presupposing a control group—'not using AI'—that no longer exists in reality. Once AI becomes infrastructure as fundamental as the internet, refraining from its use is no longer a practical option. The author argues that what truly needs evaluation are two dimensions of student interaction with AI tools: patterns of AI tool usage and the design philosophy behind those tools. The former corresponds to what many educators emphasize as part of AI literacy—active 'cognitive engagement' when using AI, avoiding 'cognitive offloading' (or outsourcing one's brain). Without distinguishing these, even identical AI tools can lead to completely opposite learning outcomes. However, merely emphasizing abstract usage patterns is insufficient. This article focuses more on the latter—the design philosophy behind the AI tools used—because this significantly influences how students develop appropriate usage patterns. Currently, the education sector's conception of AI is almost synonymous with commercial general-purpose models like ChatGPT, Gemini, or Claude, without recognizing that the entire AI industry operates by re-engineering these language models through 'harness engineering' to better suit specialized domains such as law, finance, or media. When applied to education, such specialized tools can be called 'AI tutors' or similar learning aids. The author believes a well-designed AI tutor should outperform commercial general-purpose AI across four dimensions: curriculum alignment, response guidance, student affinity, and teacher observability—enabling teachers to more effectively guide students in forming healthy AI habits, making AI literacy education far more effective, and preventing students' autonomous learning and judgment abilities from being gradually hollowed out by the flood of commercial AI. Encouraging educators to actively participate in the design, development, and selection of AI tutors naturally involves re-evaluating the policy direction and resource allocation of educational authorities in the AI era. Therefore, the answer to 'Does AI really harm student learning?' lies in whether educators have prepared suitable AI learning assistants (such as AI tutors) for students, giving them higher-quality choices during learning. Article Outline I. 'Does AI harm learning?'—A Misplaced Question II. Evidence Base: Common Structure of Two Studies III. Dimension One: Usage Patterns—Cognitive Offloading vs. Cognitive Engagement IV. Dimension Two: Tool Design—Commercial General-Purpose AI vs. Education-Specific AI Tutors V. An Ideal AI Tutor Should Excel in Four Dimensions VI. Implementation, Support Systems, and Target Students for AI Tutors VII. Conclusion: Leveraging AI Tutors to Highlight the Value of Human Teachers I. 'Does AI harm learning?'—A Misplaced Question Since the release of ChatGPT at the end of 2022, schools at all levels have responded to generative AI in roughly two phases: first, a ban period focused on academic ethics and cheating prevention; second, an encouragement phase centered on usage guidelines and workshops. In recent years, important research findings in education have been published, allowing us to assess more carefully how AI affects student learning. For example, on August 25, 2026, MIT President Sally Kornbluth issued an open letter to the entire campus and released a 40-page final report from the 'Special Committee on AI for Teaching, Learning, and Research Training,' characterizing the rapid development of generative AI as a 'watershed moment for MIT and higher education as a whole,' indicating that this technology is forcing universities to rethink the essence and operational model of education. However, the more we treat this as a watershed moment, the more we must first ask clearly: What exactly are we evaluating? The author argues that the question 'Does AI harm learning?' cannot be simply answered with a yes or no. At a deeper level, it assumes AI is a single independent variable—either fully present or absent—and directly equates it with chatbots based on large language models (LLMs), such as ChatGPT, Gemini, or Claude (hereafter referred to as commercial general-purpose AI). However, anyone with even a basic understanding of recent AI trends knows that AI applications have already entered various fields in diverse forms. The most common approach is building upon these LLMs with additional 'harness engineering' to optimize functionality for specific domain needs—examples include customer service bots in banking, legal assistants and contract review tools (Harvey AI/CoCounsel) in law, market analysis and compliance review (BloombergGPT/FinGPT) in finance, and code generation and system architecture tools (GitHub Copilot/Cursor) in software development. Unfortunately, in education, similar tools such as AI tutors designed to assist student learning are often deliberately forgotten or inadvertently overlooked in discussions about 'AI and education.' This may reflect that, compared to other sectors of society, the education field still lacks sufficient application and imagination regarding current AI technologies. Moreover, this question presupposes a control group that does not exist in reality: AI has become widely recognized as the technological foundation of the Fourth Industrial Revolution, even likened to the use of 'fire' or 'electricity'—once mastered, humanity cannot return to a pre-mastery state. This can be analogized to web search: today, libraries primarily preserve paper materials that have not yet been digitized due to the prevalence of online search. While logically we could still ask, 'Does web search harm students' ability to retrieve information?' in practice, it holds almost no meaning (unless the internet is comprehensively and long-term disrupted, which would likely only occur during war). This reflects that the education sector has not yet realized that AI has already become 'infrastructure' for future society: what can be discussed is never whether to use AI, but rather what constitutes the correct direction for its use. This article argues that what education truly needs is not an overall assessment of a flawed image of AI, but rather a distinction between two independent dimensions to evaluate AI's impact: first, students' actual patterns of AI usage, and second, the design principles of the AI platforms used. These two dimensions focus respectively on the levels of actual interaction and institutional design. More importantly, the latter greatly influences the former, elevating the discussion of 'AI's impact on learning' to the level of educational policy—a topic worthy of more careful scrutiny by scholars and experts in education. II. Evidence Base: Common Structure of Two Studies This article begins by illustrating with two representative recent research findings. The first study is a randomized controlled trial titled 'How AI Affects Skill Development,' published in January 2026 by the Anthropic research team. Participants were 52 professional software developers tasked with writing real code using a Python library they had never encountered before, with half equipped with an AI assistant. Results showed that the AI group completed tasks slightly faster than the manual group, though not statistically significant. However, in subsequent tests assessing concepts they had just coded minutes earlier, the AI group averaged 50 points, while the manual group scored 67—an 17-percentage-point gap, with an effect size equivalent to two grade levels. The largest gaps were in debugging, followed by code reading and conceptual understanding. More notably, however, the average scores among different interaction patterns within the AI group ranged from 24 to 86 points—a spread much wider than the 17-point gap between the AI and manual groups. The second study is a field experiment published in the Proceedings of the National Academy of Sciences (PNAS) by the Wharton School at the University of Pennsylvania. Researchers had nearly a thousand high school students use either a standard GPT-4 interface (GPT Base) or a teacher-designed safeguarded version (GPT Tutor) during math learning, then compared their performance in formal exams without AI. The study found that while students using the standard GPT-4 interface indeed performed worse on average than those who did not use AI, when designers simply changed the AI interface design to allow students to use the safeguarded version, it significantly deepened student-AI interaction and almost eliminated the performance gap caused by improper use of the standard interface, bringing their results close to
FACT BOX
- Source: PR Times
- Category: Survey
- Organizations: Anthropic / MIT / Pennsylvania University