OpenAI study finds ChatGPT and critical-thinking training improve different parts of student work

In a randomized experiment with more than 1,000 Bocconi students, ChatGPT improved quality and coherence while causal-reasoning training produced more varied ideas; the effects complemented each other.

OpenAI published a study on August 27, 2026 describing a randomized experiment conducted by researchers at Bocconi University in collaboration with OpenAI Economic Research. More than 1,000 first-year undergraduates worked on a marketing-recommendation case for the university merchandise store. Class periods were assigned to four groups: ChatGPT access, causal-reasoning training, both, or neither.

The results separate two outcomes that are often collapsed into one. Researchers used trained human graders and a five-point rubric, then used automated text analysis to measure the number and variety of ideas, signs of causal reasoning, and similarity to recommendations from three experts. OpenAI reports that students with ChatGPT scored almost one point higher on the five-point scale. Their submissions contained more ideas, clearer logic, and recommendations more similar to the experts' work.

That does not mean students simply handed over the assignment. The experiment still required them to decide what to ask, evaluate the responses, and choose what to include in the final submission. The tool helped narrow the visible gap between novice and expert-style output, but prompting, judgment, and selection remained part of the work.

Causal-reasoning training produced a different result. Students completed an AI-independent exercise built around a game, examples, questions, and feedback. They became better at explaining why an idea might work and when it might fail, but they did not score higher on the original rubric, which measured only two standard marketing goals: awareness and use of the university store. Automated analysis found that their ideas were more varied and less similar to their peers' answers.

The group that received both ChatGPT and the reasoning exercise showed the broadest pattern of gains. It retained ChatGPT's clearer logic and larger number of ideas while also showing more distinct ideas, hypothesis questioning, and explanations. The result does not support a simple choice between AI access and thinking skills; the two interventions addressed different dimensions of the task.

For schools and corporate training teams, the evaluation design may be the most useful lesson. If AI can help produce polished, expert-like final answers, the final answer alone reveals less about what a person understands. Assessments can ask learners to preserve their reasoning, compare alternatives, explain failure conditions, and make originality and causal reasoning explicit scoring criteria.

The study took place at one university, on one business case, with ChatGPT using GPT-4o, and OpenAI was a research collaborator. Whether the findings transfer to other ages, subjects, models, and teaching settings needs more independent and cross-context research. The defensible conclusion is narrower: in this experiment, AI helped students make answers more complete while critical-thinking training made ideas broader. Good learning design should measure both.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.