Issue #153

After AI Help, Solo Accuracy Fell Below Control Group

In one test, a group that got AI help scored 57% on the final unaided problem vs. 73% for those who worked alone throughout.

SocietyAfter AI Help, Solo Accuracy Fell Below Control Group

Can You Still Judge Alone After AI Has Helped?

Many expected that adopting AI would cut down on workload.

ActivTrak analyzed 443,000,000 hours of digital activity across more than 160,000 workers. In a sample of 10,584 people whose behavior was tracked for 180 days before and after adopting AI, time spent on email rose 104%, on messaging apps 145%, and on task-management software 94%. That said, since this is observational data comparing before-and-after periods, we can’t attribute all of the increased workload to AI alone.

Separate from the workload question is whether people who get help from AI retain their own capabilities when they later have to solve problems on their own. Recent studies suggest there can be differences in short-term performance, persistence, and how people learn.

🧠 The Group Gap That Showed Up When Solving the Last Problem Alone

A joint research team from Carnegie Mellon, Oxford, MIT, and UCLA published a paper in April 2026 reporting on three experiments involving a total of 1,222 participants. They compared groups that received AI assistance with groups that worked alone on math and reading comprehension tasks. In the first experiment, both groups solved an initial 12 problems, then tackled the final 3 problems without any AI help.

On the final problems of the first experiment, the group that had earlier received AI assistance scored 57% correct, while the control group that had worked alone from the start scored 73%. The skip rate also differed — 20% for the AI-assisted group versus 11% for the control group. This doesn’t mean the same individuals’ accuracy dropped from 73% to 57%. Rather, the result shows a gap emerged between groups when performing independently after roughly 10 to 15 minutes of assistance.

The research team likened the concern that small dependencies can accumulate to the “boiling frog effect.” But this experiment doesn’t prove long-term or irreversible skill decline. Whether a gap observed in a short task persists in education or real-world work is something that still needs further study.

A team led by Nataliya Kosmyna at the MIT Media Lab divided 54 participants into groups using ChatGPT, using a search engine, or writing with no tools at all, had them write essays, and measured their brainwaves (EEG). Brain connectivity1 measured in the ChatGPT group was lower than in the group that wrote without tools. The team also reported differences in memory and engagement, using the term “cognitive debt.” That said, an EEG difference doesn’t necessarily mean reduced intelligence or brain damage.

For an additional fourth session, some participants had their tool-use condition switched. When the group that had previously written alone then used ChatGPT, a different EEG pattern emerged compared to the group that had used AI from the start. This gives grounds to reconsider the sequence of writing first and using AI afterward, but a small, exploratory experiment shouldn’t be treated as a universal learning law that applies to everyone.

Dr. Kosmyna explained why she released a 206-page paper before formal peer review: “I was scared that in six to eight months, some policymaker would decide, ‘let’s build an AI kindergarten.’ Developing brains are the ones at greatest risk.”

Related risks have also been examined in clinical settings. Comparing AI-free colonoscopies performed before and after the routine adoption of AI at four centers in Poland, the detection rate for adenomas — precancerous lesions — fell from 28.4% to 22.4%. That’s a 6-percentage-point drop, or roughly a 20% relative decrease. This study, published in The Lancet Gastroenterology & Hepatology, is an observational study. The researchers themselves stated they couldn’t confirm AI as the definitive cause of the decline in detection rates and that further research is needed.

⚡ Getting Results Fast vs. Learning It Yourself

These results also connect to a mindset the tech industry has long followed.

David Brooks, in a recent column for The Atlantic, framed this as a split between two mindsets. Traditional education and growth rest on a “cultivation” mindset — building capability by attempting hard tasks and correcting failures. He argues the tech industry instead leans toward “optimization”: reducing user effort and delivering results fast.

Amit Singhal, former head of Google Search, said in a 2013 Guardian interview that his focus was on eliminating friction between users and information. That’s a useful goal for a search service, but education also requires the process of learning by doing. In the ActivTrak data, the group with AI usage at 7-10% showed the highest productivity metrics — and that group made up just 3% of users. This, too, is only an observed association; it doesn’t mean 7-10% is the optimal usage level for everyone.

In an experiment by a research team at ShanghaiTech University, the group that performed a task alone after receiving AI assistance showed 11% lower intrinsic motivation and 20% higher boredom than the comparison group. This suggests that an earlier experience of receiving help can carry over into attitudes toward the next task. But these figures alone don’t support the claim that everyday work becomes “unbearably tedious.”

A study from the University of Pennsylvania’s Wharton School deployed an AI deliberately designed to produce wrong answers. 80% of participants accepted the AI’s errors as-is. This result shows the risk of accepting a plausible-sounding AI answer without verifying it. Brooks calls this kind of dependence “cognitive surrender.”

One scene from Brooks’s column captures this shift vividly. A school principal in the United States showed students a film made by 200 artists over 5 years. One student asked, with a puzzled look, “Why would you do that? AI could do it in 5 minutes.” Brooks presents this question as a case study in how we evaluate the process behind a result and the effort embedded in it.

The Carnegie Mellon research team also compared, within the AI-assisted group, those who directly asked for answers versus those who asked for hints or explanations. 61% asked directly for answers, and among those who asked for hints or explanations, no clear decline in independent performance was observed. Still, this wasn’t a comparison with usage patterns randomly assigned. The people in each group may have differed in underlying ability or motivation to begin with, so we can’t conclude with certainty that asking only for hints causally preserves capability.

🇰🇷 Korea’s High AI Adoption Rate and Its Learning Environment

As AI use accelerates in Korea, we also need to look at the conditions under which people are learning while working.

According to the Bank of Korea’s 2025 AI Usage Survey, 51.8% of Korean workers use AI for work — nearly double the U.S. rate of 26.5%. Among workers who use AI, the share using it more than 1 hour a day is also higher in Korea, at 78.6%, compared to 31.8% in the U.S. In a May survey of 1,000 office workers by Now&Survey, a Korean market research firm, the share using AI at least once a week rose as high as 74.3%. The methodologies differ, but all point to the same trend: AI use is spreading fast in Korea.

What’s interesting is that the group using AI the most is also the most anxious about it. In IT and development roles, 87.6% actively use AI, and 90.7% report improved work efficiency — the highest of any occupation. Yet this same group also reported the highest sense of crisis, at 70.1%. When asked to name the “risky type” of AI user, 23.2% pointed to “people who stop thinking for themselves and depend on AI” — almost identical to the 22.8% who named “people who reject AI outright.” In other words, respondents are wary of both over-reliance on AI and blanket refusal of it.

The problem is that this awareness is hard to translate into action, given the current environment. In Microsoft’s 2026 Work Trend Index, 78% of Korean office workers said “falling behind is inevitable if you don’t adapt to AI quickly” — 13 percentage points above the global average. Yet only 16% said they felt their company’s AI direction was clear. Many respondents feel pressure to adapt without receiving enough organizational guidance on how to do so.

In PwC’s global workforce survey, only 22% of Korean respondents said they feel “it is safe to try a new approach” at work — less than half the global average of 56%. When there’s no room to experiment and fail, people naturally gravitate toward whatever gets results fastest. And the more often people simply use AI’s output as-is, the less practice they get at working out answers on their own.

Using National Pension Service enrollment data, the Bank of Korea found that youth (ages 15–29) employment fell by 211,000 jobs between July 2022 and July 2025. 98.6% of that decline occurred in industries with high AI exposure. Meanwhile, in those same industries, jobs held by workers in their 50s actually increased. The research team suggested this points to a possible “seniority-biased technological change,” in which AI replaces entry-level, routine tasks while complementing the experience of skilled veterans. That said, the entire decline in high-AI-exposure industries shouldn’t be read as jobs AI directly eliminated. What deserves closer attention is whether new hires are losing the on-the-job opportunities to build judgment in the first place.

Oswarld’s Lens

Looking at these findings, I was reminded of a pattern I kept running into during my GTM strategy consulting work.

I’ve seen companies adopt a new tool and see productivity rise at first. But once the tool starts making judgment calls instead of just supporting them, the organization’s own capacity for judgment quietly erodes. After adopting a CRM, teams lose their feel for customers. After a dashboard goes up, people stop going out into the field. Once you start relying on A/B test results, you stop asking why the results turned out that way.

You can hand AI a wide range of tasks — research, analysis, even drafting. Precisely because of that, I think we need to decide where the tool’s job ends and where our own judgment has to begin.

Brooks’ summary in this column is concise: “When intelligence is plentiful, volition is valuable.” In other words, we’re heading into an era where what determines a person’s worth isn’t intelligence itself, but the willingness to take on hard things.

To me, this isn’t just an individual problem. It’s also a problem for organizations and education systems. Right now, Korea’s workplace culture makes it hard to try new approaches. That shows up clearly in a PwC survey: only 22% of Korean respondents said they felt safe trying a new approach, less than half the global average of 56%. Alongside training people to use AI, I think we need to practice judging and verifying AI’s output for ourselves.

Closing

There’s research showing group differences in accuracy and persistence when people work on tasks alone after having used AI assistance. This doesn’t prove long-term skill decline, but it’s reason enough to design tool use and independent learning together. Looking at Korea’s high AI usage rate alongside its low tolerance for trial and error, I think it matters for organizations to build in time for people to think and verify things themselves.

I believe the people who choose to think things through on their own, even when AI is available, are the ones who end up making the difference. If you’ve ever caught yourself using AI and thought, “I used to do this myself—why don’t I anymore?”, let me know in the comments what area that was in.


📨 Please share this with someone it might help


Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Footnotes

  1. Brain connectivity: here, this refers to the relationship between signals across brain regions as measured by EEG. Its meaning varies depending on the analysis method and task, so a lower value doesn’t necessarily mean the brain “switched off” or that intelligence declined.