ChatGPT Improves Answers; Critical Thinking Widens Ideas
OpenAI published a summary of a randomized experiment with Bocconi University: more than a thousand first-year students worked the same real marketing case, randomly assigned to ChatGPT access, causal-reasoning training, both, or neither.
The finding is blunt and useful:
ChatGPT made answers better. Critical-thinking training made ideas broader. Together, the gains spanned the widest set of measures.
What the experiment measured

The assignment was a real business case: marketing recommendations for the university merchandise store—not a free-form essay.
The critical-thinking arm was not an AI prompting class. Students practiced causal reasoning—linking cause and effect, explaining why a solution might work and when it might fail—through a game, examples, questions, and feedback.
Evaluation used two lenses:
- A human five-point rubric (focused on standard marketing goals such as awareness and store use)
- Automated text analysis: idea count and variety, traces of causal reasoning, and similarity to expert recommendations
That split is the point: it separates AI access from thinking training, and also tests the combination.
AI raises “looking professional”
Students with ChatGPT (GPT‑4o) scored almost a full point higher on the five-point scale. Their work had more ideas, clearer logic, and looked more like expert recommendations.
In short: novices were lifted closer to professional-looking output.
Students still had to decide what to ask, what to keep, and what entered the final submission. AI compressed the cost of producing a coherent structure—it did not replace judgment.
For personal workflows: treat the model as a polish and expansion engine, and keep selection and decisions yours. Otherwise you only industrialize interchangeable, glossy answers.
Thinking training raises “being distinct”
The critical-thinking group had a counterintuitive result: rubric scores did not clearly rise, yet text analysis showed clearer “why it works / when it fails” explanations and a wider, more unique spread of ideas across peers.
If the rubric cannot see originality, that does not mean there was no gain. Traditional rubrics reward clear structure aimed at standard KPIs and quietly miss the cut no one else made.

That maps to a common AI-use failure mode: models pull you toward the safest, most conventional answer. Without causal chains and habit of questioning assumptions, outputs converge.
Not a trade-off—complements
The both group:
- Idea variety ≈ training-only
- Rubric scores and idea count ≈ ChatGPT-only
- Logical coherence, seeking explanations, questioning assumptions — stronger
So the tired education binary—“think for yourself” versus “learn to use AI”—is the wrong frame. Do both, and upgrade how you measure.
The real risk: grading only the final draft
Once AI makes polished, conventional answers cheap, final-answer quality tells you less about what a student actually understands.
Schools that score only “does this look like a good standard answer” will systematically reward what models can substitute and under-weight originality and reasoning process.
The same logic hits companies and indie builders:
- Add evaluation dimensions: process assumptions, failure conditions, differentiated ideas—not just verbal polish.
- Don’t only design human tabs: interaction may converge into an AI layer, but capabilities still need to be callable and checkable.
- Personal stack: causal reasoning plus saying things clearly remains the hard leverage—ask well, filter well, and the final draft is yours.
One line to keep
AI helped students make their answers better. Critical-thinking training helped make their ideas broader.
Tools raise the waterline; training widens the map. Miss either side and you get homogeny or roughness. Practice both—that is the preparation this experiment actually argues for.
Source: OpenAI — What students gain from ChatGPT and critical-thinking training
Comments