AI Products & PlatformsCommunity & Cognition

AI Raised the Floor. The Grading Standard Still Punishes the Ceiling.

A Phony War

Whenever tech circles discuss the impact of AI on learning, people usually split into two camps quickly. One camp views the chat box as a 24/7 tutor—always ready to answer questions, letting you keep moving whenever you get stuck. The other camp sees it as nothing more than glorified homework copying: throw the problem into the dialogue box, grab the ready-made answer, and over time your thinking habits wither. I ran into both camps when discussing where AI education should land four months ago: the prevailing narrative back then was to pair every student with an AI tutor, yet the Turkish high school experiment had already hinted that the exact same model produces opposite outcomes under different usage modes.

Two thousand-student randomized controlled trials in academia have now brought these two polar-opposite trajectories out in plain sight. The first took place in Milan, published by OpenAI’s economic research team in collaboration with a university, focusing on a business marketing assignment. 1053 freshmen were asked to produce an 180-word proposal within 45 minutes to help their alma mater’s souvenir shop boost alumni awareness and purchase intent. Students with access to ChatGPT scored nearly a full point higher on a 5-point scale than the control group, with strategies and phrasing visibly approaching the reference answers handwritten on site by three experts. Looking at the submissions alone, beginners’ work unmistakably approached expert-level reference solutions.

The second was a high school math experiment in Turkey, published formally in PNAS. Across roughly 50 classes, nearly 1000 high school students went through 4 sessions of 90-minute instruction and practice, with all scores counted toward their final grades. During the in-class practice phase, students assisted by standard ChatGPT scored 48% higher than the control group. Yet the twist arrived immediately. As soon as practice ended, students took a closed-book, screens-off exam within the exact same class period. The very students who had just led practice by 48% saw their independent exam scores drop by -17% compared to the control group. Away from the screen, they immediately fell behind classmates who had relied solely on textbooks and notes.

The same tool pushed expression close to expert level in a business proposal, yet caused students’ test scores to drop by nearly twenty percent once away from the screen in a math exam. These seemingly contradictory trajectories are actually a phony war, born from a mismatch in measurement dimensions. To see the dividing line clearly, we only need to unpack three terms: the floor, the ceiling, and the grading standard. Writing smooth and complete copy, structuring cleanly, and eliminating glaring flaws represents the floor; probing underlying assumptions, tracing causal chains, and identifying the boundaries where a proposal fails represents the ceiling. And what ultimately determines a student’s grade is what the grading standard behind it actually rewards. Taking these three terms back to the ground, the seemingly conflicting conclusions reveal a lucid logic.

The Math Experiment: The Floor Is Borrowed

In the Turkish high school math classrooms, students tackled 57 problems in algebra and geometry. Each 90-minute class was split into three segments: lecture review, assisted practice, and unassisted testing. During the practice phase, screens lit up, and AI answered questions on demand. The control group had only textbooks and notes, averaging 0.28 during practice. Students assisted by standard ChatGPT saw their scores increase by 0.137, a relative gain of 48%; the tutor mode group, configured with strict rules, saw scores increase by 0.361—a relative surge of 127%. Looking strictly at the page at this stage, tutor mode more than doubled problem-solving performance.

The real test came immediately afterward in the closed-book exam. As soon as practice ended, screens went dark, and students received conceptually similar questions to answer closed-book. The control group averaged 0.32 at this stage. The standard ChatGPT group saw test scores drop by 0.054, dragging exam performance down by -17%. Meanwhile, the tutor mode group—which had surged 127% during practice—experienced an exam score change of just -0.004, statistically indistinguishable from the control group. The dazzling advantage on screen vanished the moment the network was cut: the harm reset to zero, and the gains reset to zero just the same.

Solving problems in front of a screen felt effortless, but once away from the tool, students could not solve problems independently. Researchers analyzed the model’s performance across the 57 problems and found that standard ChatGPT produced the correct answer only 51% of the time on average, with reasoning errors accounting for 42% and arithmetic errors accounting for 8%. Had students followed along with the reasoning steps, these logic flaws should have interfered with the subsequent exam; yet empirical tests showed that this bias did not spill over into subsequent testing, with a correlation coefficient of only -0.029, which is statistically insignificant. Students were not misled by flawed logic simply because they never read the derivations. Backend logs revealed that the students’ most common action was asking directly for the final answer and copying it into their workbooks to secure high marks. Students bypassed inferential thinking entirely, reducing practice to moving answers from the screen onto paper.

To block students from requesting ready-made answers, the research team recruited two part-time math teachers to compile standard solution paths and common error libraries problem by problem, configuring the system with a strict tutor mode: it was forbidden from giving direct solutions, mandatory for students to first show which step they had reached and where they were stuck, and programmed to dispense clues only incrementally at minimal increments, while verifying student-provided answers. This tutor mode curbed the impulse to copy final answers and prevented test scores from regressing. Yet even with this labor-intensive human configuration, when students answered questions independently away from the tool, their performance still failed to surpass ordinary students who only had textbooks and notes. Post-class self-assessments revealed a striking dissonance: the standard ChatGPT group performed worse yet felt great about themselves; the tutor mode group showed no gain in exam scores, yet gave significantly higher subjective ratings. A seamless interaction easily manufactures the illusion of competence.

In the Turkish high school math experiment, student practice scores rose by nearly half or even doubled when AI was present, but closed-book exam scores dropped rather than rose once AI was removed.

The contrast between in-class practice and independent testing points to the first yardstick for evaluating technological impact: whether AI is actually present during measurement. Reviewing empirical studies over the past two years reveals a stark divide: every finding that technology impairs learning comes from unassisted measurement after AI is removed; virtually every report celebrating dramatic leaps in learning records performance with AI present or under real-time assistance. While the screen is glowing, the test only measures a human-machine composite. The tool can elevate the floor of an assignment to a respectable standard at negligible cost, but this floor is borrowed from external compute. Once you unplug the cable and shut off the screen, the borrowed floor vanishes in an instant, leaving behind a brain that never experienced cognitive friction facing the test paper.

The Marketing Assignment Experiment: The Grading Standard Only Recognizes the Floor

Math problems have objective right and wrong answers. In business pitches and open-ended humanities tasks without standard solutions, however, things become far more subtle. The marketing assignment experiment published by OpenAI’s economic research team in collaboration with a university turned the lens toward business writing. 1053 freshmen were asked to write a marketing proposal of no more than 180-word within 45 minutes for their alma mater’s souvenir shop to enhance alumni awareness and purchase intent. 20 master’s students served as independent blind graders, with each submission independently evaluated by 3 raters, totaling 3159 evaluations scored on a 1-to-5 scale.

The introduction of the tool quickly demonstrated its power to raise the floor. The student group using ChatGPT scored +0.862 points higher than the control group. The baseline score for the control group was just 2.09 points; with model assistance, the students’ average score jumped to 2.95 points (calculated with rounding). On a 5-point scale, an advantage of nearly 1 point is substantial. Three experts—a marketing professor, a souvenir store manager, and an alumni specialist—handwrote reference proposals averaging 208-word on site to eliminate pre-training data contamination. Textual analysis revealed that the ChatGPT group’s similarity to expert reference solutions increased significantly across all three metrics. Technology genuinely helped inexperienced freshmen leap over expressive hurdles to produce seasoned proposals.

Yet what followed took an unexpected turn. The research team designed an additional intervention: a thinking workshop lasting just 9 minutes, featuring a minigame with 12 questions before writing began, nudging students to focus on two things: deriving the causal mechanisms behind why a strategy works, and identifying under what conditions a strategy fails. In just 9 minutes, the way students approached thinking in this assignment shifted. Textual analysis showed that trained students exhibited markers of deeper cognitive engagement in their submissions: they were better at articulating the causal mechanisms that made strategies work, and significantly more proactive in stating under what conditions their proposals would fail. A proposal that not only offers strategies but also unpacks mechanisms and marks failure boundaries ought to receive higher praise. Yet the grading outcome was chilly: the group that received only the thinking training saw their score dip slightly by 0.111 points; combining thinking training with ChatGPT yielded scores indistinguishable from using ChatGPT alone. Spending effort to think deeply earned no reward on paper.

Why did thinking deeper fail to yield higher scores? The research team dissected the rubric line by line, and the figures revealed its true orientation. Using statistical models to attribute grading outcomes to specific features, researchers found that submissions explicitly stating failure conditions received systematically lower scores (a negative association of 0.11 to 0.16 points); answers explaining mechanisms were likewise penalized (0.17 to 0.24 points); and diverging from conventional solution templates was similarly docked. Conversely, answers with smooth, coherent logic scored higher (associated with 0.355 to 0.547 points), and submitting a higher number of ideas boosted scores (adding 0.09 to 0.12 points for each additional idea). Within the 180-word limit, conscientious students spent precious word count defining assumptions, explaining mechanisms, and highlighting risks, which graders perceived as lacking decisiveness. Meanwhile, students who effortlessly rattled off a few promotional formulas sailed through with top marks. The rubric’s yardstick measured only the completeness of boilerplate formulas; rigorous inquiry exploring the cognitive ceiling became an unwelcome penalty.

This finding overturned my initial assumption. When I first started tracking this study, my initial reaction was that AI was funneling everyone’s thinking into homogenized mediocrity; I noted at the time that generative models acted as formula generators, flattening unique ideas and pouring everyone into the same mold. But after reading through the paper’s full dataset, I realized I had reversed the attribution. The data clearly shows that the distance between different ideas in the ChatGPT group did not shrink: the statistical coefficient for between-solution variance was -0.002, which is statistically insignificant. AI did not compress ideas into a narrow mold, nor did it flatten the diversity of thought.

What exerted the standardizing pressure was the entrenched, conventional rubric itself. This rubric favored well-rounded, predictable standard answers, while digging into premises and exposing limitations only invited point deductions. And when it comes to drafting standard proposals cleanly and coherently, AI has made that virtually effortless. AI acts like a mirror, exposing how traditional grading standards have long suppressed the cognitive ceiling. As the paper notes at the end of its abstract, for diverse ideas to realize their value, evaluation systems must actively encourage divergence and exploration rather than merely rewarding standard solutions.

In the Milan business school experiment, detailing failure conditions and operational mechanisms earned students no extra points, and the rubric penalized such content instead.

This brings us to the second yardstick: what the grading standard in place actually rewards. When an assignment secures a stellar score with the help of AI, there is no need to celebrate prematurely. We need to examine the yardstick held by the evaluators to see whether it encourages the ceiling of exploring the essence of a problem, or merely dishes out rewards to the floor that caters to conventional formulas.

How to Put This to Work

With these two yardsticks in view, how we should live with AI in our daily work and study becomes quite concrete, boiling down to three actions. First, use AI with confidence when it is present; the floor has genuinely been raised. Whether drafting documents, compiling meeting summaries, or writing routine code, the model’s ability to frame structures and flesh out details is undeniable. Relying on tools to shore up the baseline floor and free humans from tedious, entry-level drafting is, in itself, tangible productivity.

Second, regularly pull AI away to test yourself and verify whether the capability truly belongs to you. A polished deliverable cannot prove that its author has mastered the underlying logic. Looking at the final output alone, you cannot tell whether someone has genuinely learned or simply rented a respectable output from an algorithm. Both teams and individuals need to establish unassisted, closed-book scenarios—deriving core logic on a blank page to confirm that critical capabilities reside in their own minds.

Third, revise the grading standard so that it can see the ceiling. If an evaluation system only hands top marks to fluent prose, tidy formatting, and conventional tropes—or worse, penalizes deep thinking that questions boundaries, uncovers risks, and investigates mechanisms—everyone will inevitably drift into using models to mass-produce mediocrity. To steer toward deeper inquiry, we must overhaul the direction of evaluation, shifting the focus from unilaterally rewarding standard answers to actively encouraging the testing of assumptions, the derivation of causal mechanisms, and the exploration of non-standard solutions.

When examining the impact of technology on learning, we also need to maintain lasting intellectual honesty and restraint. To date, not a single randomized controlled trial has proven that an AI configured in tutor mode with strict rules produces positive gains on independent exams once the tool is withdrawn; as for long-term retention, such trials have not even tested it yet. In the Turkish high school math experiment, even with two part-time math teachers painstakingly cataloging standard solutions problem by problem, students’ performance on closed-book exams merely matched that of the control group.

The most robust positive evidence found in the real world today comes from technology assisting human instructors. In a randomized controlled trial covering 700+ tutors across math coaching programs in disadvantaged US communities, when tutors were equipped with real-time expert recommendations, student mastery rates increased overall by 4 percentage points; among tutors who initially had lower ratings, mastery rates jumped significantly by 9 percentage points. Only when technology was deployed to help human tutors pinpoint bottlenecks and optimize interactions did it deliver dependable gains. In this direction, I proposed a hypothesis in an earlier article analyzing leverage points in AI education: the larger variable determining educational outcomes likely lies in lesson design quality rather than whether each student receives individualized attention; AI’s greater leverage may not lie in assigning an AI teacher to every student, but in helping teachers make classes better. Today, these two experiments provide fresh evidence for that assessment: even the most successful application of ‘providing an AI teacher’ works by putting AI in the instructor’s corner.

When the grading standard is only willing to pay for an unremarkable floor, even the most powerful tools will only churn out cheaper formulas; only when the evaluation system begins to reward the ceiling that breaks away from convention will the human brain truly start training its own thinking. In the final analysis, both experiments yield the exact same answer: what you evaluate determines what you train.

Sources