A new study of more than 1,000 university students offers one of the clearest empirical answers yet to a question that has hung over classrooms since ChatGPT arrived: does the tool help students learn, or does it just help them produce polished work without actually thinking?
The answer, according to a randomized controlled trial published by OpenAI on August 27, 2026, is: it depends what you’re measuring. Access to ChatGPT made student work noticeably better by conventional grading standards. But it did nothing to make that work more original. That benefit came from something else entirely — a short training exercise in causal reasoning that had no connection to AI at all.
The two interventions, tested side by side in the same experiment, turned out to do different jobs. Neither substituted for the other.
Table of Contents
What Happened
The study was conducted by researchers primarily affiliated with Bocconi University in Milan, in collaboration with OpenAI’s economic research team, and was published as a working paper alongside OpenAI’s announcement. The research team included Bocconi and SDA-Bocconi academics, one OpenAI researcher who also holds a Duke University affiliation, and two additional contributors — one who worked on the study while employed at OpenAI and has since left, and one who worked as a paid OpenAI contractor while affiliated with the University of California, Berkeley.
Notably, the working paper itself withholds the name of the university involved, citing a need to preserve author anonymity tied to its pre-registration with the American Economic Association’s RCT registry. OpenAI’s own blog post, along with the authors’ listed institutional affiliations, identifies the institution as Bocconi. That inconsistency between the paper’s stated anonymity and the company’s public naming remains unresolved.
The trial itself was conducted in November 2025, though results weren’t made public until several months later.
Timeline
- November 2025: Data collection takes place over two consecutive days at Bocconi University.
- August 27, 2026: OpenAI publishes its blog post and the underlying working paper.
- Now: No independent replication, peer review, or third-party commentary has yet been published.
How the Study Worked
Researchers recruited 1,053 first-year undergraduates across 13 sections of an introductory management course, spanning three degree programs in economics, finance, and management. Students were randomly assigned, by class section, to one of four groups: no intervention, access to ChatGPT (specifically GPT-4o, through an institutional “ChatGPT Edu” subscription), a short training exercise in causal reasoning, or both combined. The combined group was deliberately oversampled to give researchers more statistical power to detect interaction effects.
The causal-reasoning training was not framed to students as AI literacy or general “critical thinking” instruction in the abstract. It was a structured, gamified exercise — built around a hypothetical city-policy scenario — that walked participants through three specific skills: building coherent cause-and-effect chains, stating falsifiable predictions, and identifying underlying mechanisms behind an outcome. Researchers avoided using the term “causal reasoning” with participants, describing it instead as practicing “critical thinking skills.”
After the training or placebo game, all students completed the same real assignment: a 45-minute business case asking them to write roughly 180 words of consultant-style recommendations for increasing alumni engagement with Bocconi’s own branded merchandise shop. Students in the ChatGPT groups could use the tool in a separate browser tab; those without access completed the task unaided.
Twenty trained master’s students, working in teams of three per submission, graded each response on a five-point rubric measuring two standard marketing concepts: awareness and usage. Separately, researchers compared each student’s answer to three professionally written expert solutions — from a marketing professor, a merchandising shop manager, and an alumnus with relevant expertise — none of whom were permitted to use AI while writing their own reference answers.
At a glance: the four experimental groups
| Group | Students (N) | Rubric score effect | Idea diversity effect |
|---|---|---|---|
| Control | 249 | — (baseline) | — (baseline) |
| Critical-thinking training only | 256 | Slightly negative | Increased |
| ChatGPT access only | 197 | Increased (~0.86 pts) | Little change |
| Both combined | 351 | Increased (matched ChatGPT-only) | Increased (matched training-only) |
What ChatGPT Access Did
The effect of ChatGPT access was large and consistent. Students given access to the tool scored roughly 0.86 points higher on the five-point rubric than the control group, whose average score sat at about 2.09 to 2.24 depending on the model specification — a substantial jump on a scale that only runs from one to five. Their answers were also measurably closer, in semantic terms, to the professionally written expert solutions.
Part of that advantage came down to straightforward mechanics: ChatGPT-assisted answers contained more distinct ideas and read with clearer logical structure. But the researchers also tested whether the effect was purely cosmetic — a matter of better writing rather than better thinking — by statistically controlling for text complexity, sentence structure, and the number of ideas present. Even after those controls, roughly a third to half of ChatGPT’s original advantage remained, suggesting the tool was improving the substance of students’ recommendations, not just their polish.
Critical-thinking training alone, by contrast, did not raise rubric scores. If anything, the effect was slightly negative, though not dramatically so.
What Critical-Thinking Training Did
Where the training showed its value was in a dimension the rubric never measured: originality.
Researchers built a separate pipeline to assess how varied and distinct each student’s ideas were, both within their own submission and compared to the ideas of every other student in the sample. This wasn’t done by human judges reading for creativity — it was computed. A separate AI model (GPT-5.2) extracted the individual ideas embedded in each response across multiple runs, a second layer of review reconciled inconsistencies between runs (using both GPT-5.2 and Anthropic’s Claude Opus 4.6 as cross-checking judges), and the ideas were then converted into embeddings and compared mathematically using cosine distance — a standard way of measuring how semantically different two pieces of text are from one another.
By that measure, students who completed the causal-reasoning training produced ideas that were both more varied within their own answer and more distinct from what their classmates wrote. ChatGPT access, on its own, did not move this measure much. The training was also the only intervention that meaningfully increased students’ use of falsifiable statements and explicit cause-and-mechanism reasoning — the specific skills it was designed to teach.
The researchers’ explanation for the rubric gap is direct: the standard marketing rubric used in the experiment rewarded conventional, well-structured answers and did not give credit for — and in some cases statistically penalized — ideas that departed from familiar solutions. In other words, the training worked. The measurement just wasn’t built to notice.
Combining Both
Students who received both interventions captured much of what each one offered individually. Their idea diversity roughly matched students who received training alone, and their rubric scores and idea counts were comparable to students who had ChatGPT alone. This group also showed the strongest gains in logical coherence and reasoning quality across the widest range of measures the researchers tracked.
That said, the combination wasn’t a straightforward multiplier. Several of the statistical interaction terms — measuring whether pairing the two treatments produced something beyond what each contributed separately — were small and, in some analyses, not statistically distinguishable from zero. The clearest way to describe the combined effect is additive and complementary, not synergistic in a dramatic sense.
What This Means for Schools and Assessment
The study’s authors draw a specific and somewhat pointed conclusion from their own data: if AI can help students produce polished, expert-resembling work regardless of whether they’ve developed the underlying reasoning skills, then grading systems built around rewarding conventional, well-structured answers are measuring less than they used to. A rubric that penalizes distance from the typical answer, they argue, will systematically undervalue exactly the kind of original thinking that critical-thinking instruction is meant to cultivate.
That is the researchers’ interpretation of their own findings, not an independently adjudicated verdict on how universities should grade assignments. But it does align with a broader concern raised elsewhere in AI-and-education research: that as AI narrows the gap between novice and expert output, the qualities that once signaled genuine understanding — a well-organized, comprehensive answer — become easier to produce without much of that understanding at all.
How This Fits Broader Research
This study sits alongside — and doesn’t entirely resolve — a mixed body of existing evidence on AI in education.
A 2025 study led by researchers at the University of Pennsylvania’s Wharton School examined generative AI tutoring in a very different setting: high school math classrooms in Turkey. That research found something more troubling than the Bocconi study’s results — students who used an unrestricted AI tutor performed markedly better on practice problems while assisted, but performed worse on a later unassisted exam than students who’d had no AI access at all, suggesting the tool functioned as a crutch that undercut retention. The lead researcher on that study described the pattern as students not realizing they were using AI to avoid doing the cognitive work themselves.
Separately, MIT’s Media Lab published widely covered research finding that participants writing essays with ChatGPT showed reduced brain connectivity and weaker recall of what they’d written, compared to participants using a search engine or no tools at all.
Neither of those studies measured the same thing the new Bocconi-OpenAI research measured. The Wharton study tracked learning retention over a longer arc; the MIT study measured neural engagement during writing. The new study measured the quality and originality of a single piece of work completed in one sitting. Together, the three studies point to a more complicated picture than any single headline suggests: AI can improve what a student produces in the moment while leaving open real questions about what they retain, and whether measuring a finished product is the same as measuring whether they learned anything.
Limitations and Open Questions
Several caveats are worth keeping in mind before drawing broad conclusions.
The study was conducted at a single university, using a single assignment, with a specific version of ChatGPT — GPT-4o, not a current-generation model. Its findings may not generalize to other tasks, disciplines, or newer AI systems. The measure of “performance” was tied to one particular grading rubric built around two marketing concepts; a different rubric, designed to reward original thinking, might have produced different results entirely.
The research also measured only a single 45-minute task, not what students retained afterward or whether the reasoning skills the training conveyed would hold up over a semester or a career. The authors are candid about this limit themselves, noting they cannot determine whether ChatGPT’s advantage reflected knowledge students actually absorbed, or content they simply obtained without learning anything.
It’s also worth noting who conducted the research. While most of the authors are academic researchers unaffiliated with any AI company, the author list includes people who currently work at OpenAI, previously worked at OpenAI, or were paid as OpenAI contractors — a detail relevant to how independently the results should be weighed. And as mentioned earlier, the paper’s own anonymization of the university’s identity, despite OpenAI’s public post naming it, remains an unexplained inconsistency. Finally, there’s no indication yet that the paper has undergone or is undergoing formal peer review; it appears, at least for now, as a working paper distributed directly through OpenAI’s own channels.
Conclusion
The evidence here is specific and, within its scope, fairly solid: in this one experiment, ChatGPT closed much of the gap between novice and expert output on a well-defined task, while a short course in causal reasoning made students’ thinking more original in ways a conventional grading rubric didn’t capture — and didn’t disappear when ChatGPT was available too.
What remains uncertain is how far any of this travels beyond the specific classroom, task, and AI model tested. Whether the same pattern holds for different assignments, older or newer AI systems, or longer time horizons is not something this study can answer. Readers should also treat “originality,” as measured here, for what it is — a computational estimate of semantic distance between texts, not a human judgment of creative merit.
What’s worth watching next is whether independent researchers examine or replicate these findings, whether the study eventually appears in a peer-reviewed journal, and whether universities begin experimenting with grading rubrics explicitly designed to reward original thinking rather than penalize it.
Frequently Asked Questions
What did the new OpenAI-backed ChatGPT study find?
It found that giving university students access to ChatGPT improved the quality of a real assignment — raising rubric scores and making answers more similar to expert solutions — while a separate critical-thinking training exercise improved the originality and diversity of students’ ideas without raising rubric scores.
Which university ran the study, and who conducted it?
OpenAI’s announcement and the researchers’ listed affiliations identify Bocconi University in Milan as the site of the study, conducted with OpenAI’s economic research team. The working paper itself withholds the university’s name, citing anonymity tied to its trial registration — an inconsistency that hasn’t been publicly explained.
Did ChatGPT make students’ work more original?
No. ChatGPT access improved rubric scores and made answers resemble expert solutions more closely, but it had little effect on the diversity or originality of ideas. That benefit came specifically from the causal-reasoning training.
Does critical-thinking training help if students can use ChatGPT anyway?
Yes, according to the study. Students who received both the training and ChatGPT access retained most of the originality benefits of training and most of the quality benefits of ChatGPT — the two effects didn’t cancel each other out.
Was this study peer-reviewed?
There’s no indication that it has been. It was published as a working paper directly through OpenAI’s own channels, without a stated journal affiliation or peer-review status.
How is “originality” measured in the study?
Researchers used AI models to extract the distinct ideas in each student’s response, then measured how semantically different those ideas were from each other and from other students’ ideas, using embedding-based distance calculations rather than human judgment of creativity.
How does this compare to other research on AI and student learning, like the Wharton or MIT studies?
Those studies looked at different things. A 2025 Wharton-led study found unrestricted AI tutoring improved practice performance but hurt unassisted exam scores in high school math, pointing to retention concerns. MIT’s Media Lab found reduced brain engagement when writing essays with ChatGPT. This new study measured the quality and originality of a single assignment, not long-term retention or neural activity, so its results complement rather than directly confirm or contradict those findings.

