467 Agreed to Quit Phone Internet. 119 Actually Did.
A two-week phone-internet detox study reveals more in its 25.5% compliance rate than in its headline mental-health results.
AI & Tech467 Agreed to Quit. 119 Actually Did.
Let me tell you about an experiment.
A joint research team from Canada and the US made an offer to 467 people: “Would you be willing to completely cut off mobile internet on your smartphone for two weeks?” They could still make calls and send texts, and use the internet on their laptop. Only the phone itself would go offline. The compensation was generous, too.
467 people said yes. And 119 people actually made it through the full two weeks.
This study1, published in PNAS Nexus in February 2025, was mostly reported with the takeaway that “cutting phone internet improved mental health.” That summary is accurate. But the number I lingered over longest in this paper wasn’t the effect size — it was that 25.5% compliance rate. What’s worth noticing here is that even people who already believed phone internet was harmful and deliberately decided to quit it mostly couldn’t keep it up for two weeks.
Study Design, Screen Time, and Changes Across Three Metrics
Let’s start with how the experiment actually ran.
Participants installed an app called Freedom. Its “lockdown mode” makes it impossible to disable the block from within the app itself. The research team used this app to objectively track whether the block was actually turned on — they didn’t rely on participants’ self-reports. That’s a real methodological strength of this study.
The intervention group’s average daily screen time dropped from 314 minutes to 161 minutes. That’s 5 hours 14 minutes down to 2 hours 41 minutes. The three outcome measures moved as follows:
- Subjective well-being (life satisfaction + positive/negative affect): d = 0.452
- Mental health (depression, anxiety, anger, social anxiety, etc., measured on American Psychiatric Association scales): d = 0.56
- Sustained attention (objectively measured via gradCPT3): d = 0.23
90.7% of participants showed improvement in at least one of the three. The research team described the improvement in attention as “equivalent to reversing a decade’s worth of age-related decline.” That’s a striking figure for a two-week intervention.
There’s one claim in this paper worth double-checking before you cite it. The most widely quoted line is that “the improvement in depressive symptoms exceeds the meta-analytic effect size of antidepressants.” The basis for this comparison is the 2008 Kirsch et al. meta-analysis4, which re-analyzed FDA submission data and sits on the low end of estimates for antidepressant efficacy — and remains contested to this day. The effect-size estimate of 0.32 itself has been fought over, re-analyzed, and rebutted for well over a decade now. Swap out the comparison point, and this claim may no longer hold. Worth keeping in mind if you plan to cite it.
The stage where 467 agreers dropped to 119
Here’s how 467 people became 119, step by step.
| Stage | Count | Note |
|---|---|---|
| Agreed to 2-week block, received code | 467 | |
| Actually registered the app code | 272 | 195 didn’t even sign up |
| Set up initial block | 266 | |
| Completed all 3 surveys | 249 | |
| Kept the block on 10+ of 14 days | 119 | 25.5% of agreers |
First, 195 people—41.7%—didn’t even sign up for the app. They’d agreed to be paid, said yes, even received the code, and then just… didn’t do it. And of the 249 who installed the app and answered every survey to the end, just under half—119 people (47.8%)—actually maintained the block.
Here’s an important piece of context: these participants weren’t randomly selected members of the public. At the start of the experiment, 83% rated their motivation to cut phone use at 5 or higher out of 7, and 79% said they had sufficient ability to cut back. In other words, these were people who were both motivated and confident.
25.5% of those people lasted the two weeks.
One more thing. After the intervention ended, screen time bounced back from 161 minutes to 265 minutes. It didn’t fully return to T1’s 314 minutes, but once the enforcement disappeared, most people drifted back to roughly where they started—even people who had physically felt the improvement.
It’s hard to explain this result as individual weakness. If the same outcome showed up for the vast majority of people who said they wanted to cut back and believed they could, then the cause is more likely to lie in the environment surrounding the phone and its apps than in anyone’s personal willpower.
The Paradox of a Sense of Self-Control
The research team used mediation analysis5 to figure out why the improvement happened. The results are interesting.
Here’s the order of magnitude for the pathways explaining the improvement in subjective well-being.
- Increased sense of self-control (0.169) ← overwhelmingly the top factor
- Increased social connectedness (0.084)
- Increased offline time (0.072)
- Decreased media consumption (0.046)
- Increased sleep (0.020)
Once you strip out all five of these pathways, the direct effect of the intervention statistically disappears (c′ = 0.034, not significant). For mental health, too, sense of self-control (0.209) ranked first. And in reality, sense of self-control shifted by d = 0.66 between before and after the intervention — the second-largest change among the five mediators, after offline time (d = 0.70).
What people felt they’d regained was “a sense of self-control” — but what they’d actually done was give up self-control and hand it over to an app.
Freedom’s lock mode isn’t a feature that strengthens willpower. It’s a feature that eliminates the very opportunity to exercise willpower in the first place. Participants didn’t win a moment-by-moment battle over “should I open Instagram right now or not” — they entered an environment where that question never even arose. And yet the sensation that resulted was “I am in control of my life.”
A sense of self-control, it turns out, wasn’t the cause of self-control — it was the byproduct of good environmental design.
One more thing worth noting. Improvement in attention alone was not explained by any of these five pathways. Offline time, sleep, sense of self-control — none of them mediated the improvement in attention. This means well-being and attention are damaged and restored through entirely different channels. Nobody knows the answer to this yet, and I think this gap is the most honest part of the paper.
Oswarld’s Lens
There’s a pattern I keep running into while building GTM strategy.
Features that depend on user willpower fail almost without exception. The moment onboarding asks users to “go into settings and turn this on,” conversion collapses. Features that encourage “build a habit” leave no trace on the retention graph. Conversely, I’ve watched a single default-value change move an entire metric overnight. In product circles this is already common sense: defaults shape behavior far more than user intent does.
But we rarely apply this common sense outside the product itself.
The entire digital wellbeing industry is currently designed in exactly the opposite direction. Screen-time dashboards show you statistics, induce a bit of shame, and then hand the decision back to you. Usage alerts inform you that “you’ve exceeded your set limit” — and then place a “Dismiss” button right next to the message. What every one of these features has in common is that the decision to actually cut back is ultimately left to the individual user’s willpower. The 25.5% compliance rate in this paper is exactly the data point showing what happens when that decision gets left to the individual.
This bothers me, because we’re the ones designing that environment. When engagement is the KPI, what we end up building is a service that users later have to block with a separate app. The fact that apps like Freedom sell at all reveals the underlying structure: services designed to maximize usage time create a problem, and a paid app that blocks usage solves it for a fee.
And right now, that loop is tightening even further. Most AI assistant product narratives lean toward being “always present, always aware of context, always the one to speak first.” That runs directly counter to what this paper implies. To be clear, it would be premature to claim AI assistants are equivalent to a barrage of notifications. But I do want to flag this: right now we keep asking “how much more present can we make it,” and almost never ask “when should it disappear.” Not showing up in front of the user when it isn’t needed should also count as a feature a product needs to have.
Three Limits: Sample, Expectation Effects, and Conflicting Research
Let me be honest about the limits, too.
The sample was unusual. Only iPhone users, drawn from Prolific, an online research labor pool, averaging 32 years old, 63% women. What’s more, 83% were people who already wanted to cut back on their phone use. Put differently, nobody knows whether the same effect would show up in someone who has no interest in cutting back.
Expectation effects weren’t ruled out. Participants knew this was a study on “the effects of smartphones on well-being.” There may have been pressure to report feeling better. The research team flagged this limitation themselves, which is exactly why they included gradCPT instead of relying on self-report alone. But here’s what’s interesting: the effect on the objective measure, attention, was the smallest (d = 0.23). The pattern — bigger effects on the more self-reported measures — is worth paying attention to.
There’s also conflicting evidence out there. Orben and Przybylski, using large-scale time-diary data, reported that screen time explains only about 0.4% of the variance in adolescent well-being. Granted, this analysis itself has drawn criticism for “analytical choices that shrank the effect,” so it’s hard to say definitively which side is right. But the gap itself — small effects in correlational studies, large effects in experimental ones — is a signal that we still don’t know how to measure this phenomenon properly.
Even accounting for all three of these limits, something remains. A 25.5% compliance rate isn’t explained away by expectation effects or sample bias. If anything, the number stands out more precisely because the sample skewed toward the motivated. Even among people who said they wanted to cut back, 3 out of 4 couldn’t hold to it for 2 weeks.
Closing
Let me sum up.
- Cutting off phone internet for 2 weeks genuinely improves well-being, mental health, and attention. That said, you should cite the effect-size comparisons carefully.
- The single biggest driver of improvement was a sense of self-control — yet what participants actually did was hand self-control over to an app. A sense of self-control isn’t a product of willpower; it’s a product of environmental design.
- So the practical takeaway from this study isn’t “make a resolution” — it’s “build an environment where you don’t need one.” For what it’s worth, the research team itself doesn’t think a total block is necessary. The effect showed up even in the ITT analysis6, which included everyone who failed to stick to the block — meaning simply cutting back may be enough.
Let me add one point about the Korean context. In the 2025 Smartphone Overdependence Survey, the overall at-risk group7 fell to 22.7%, its fifth consecutive year of decline — but among teenagers alone, it actually rose to 43.0%. When everyone else is improving and one group is moving in the opposite direction, that’s more likely a problem with that group’s environment than with their willpower. This is a place where the paper’s framework is worth applying to the smartphone environment teenagers are living in.
Here’s something you can try right now. Instead of deleting apps or setting time limits, try creating just one situation where you have zero opportunity to exercise willpower at all. Not bringing your phone into the bedroom works. Turning off data on your commute works too. The point isn’t “deciding not to look” — it’s designing a state in which you can’t look. In this study, only 119 of the 467 people who initially agreed to the experiment actually stuck with the plan to the end. We need to keep in mind just how hard it is to maintain a resolution to cut back once you’re back in ordinary life.
If you’ve ever tried to cut back on phone use and failed, tell me in the comments exactly where it broke down. Was it a notification? A moment when your hand reached for it out of habit? Or something that made you anxious if you didn’t check it? That’s precisely the question this paper leaves unanswered — so your experience could become material for the next issue.
💬 Where did it break down, that moment you failed to quit your phone? · 📨 Share this with someone it might help
(Web version) 📬 Subscribe · I write a weekly newsletter that cross-analyzes technology, economics, and the humanities.
Keep the perspective, not the noise.
We choose one consequential shift and trace what sits beneath it, every other day.
Confirm once to finish subscribing.
Already a subscriber? Sign in to join the conversation
References & Further Reading
Primary sources
-
Castelo, N., Kushlev, K., Ward, A. F., Esterman, M., & Reiner, P. B. (2025). “Blocking mobile internet on smartphones improves sustained attention, mental health, and subjective well-being.” PNAS Nexus, 4(2), pgaf017. Read the paper ··· This is today’s core evidence. I’d recommend looking at the CONSORT diagram in “Materials and methods” (Fig. 3) before the results section — it shows the entire journey from 467 participants down to 119 on a single page.
-
Raw data and preregistration for this study: OSF repository ··· You can check for yourself exactly where the preregistration and the actual analysis diverged. The paper itself notes that the time-use factor analysis was a post hoc analysis, not a preregistered item.
Background
-
Kirsch, I. et al. (2008). “Initial severity and antidepressant benefits: a meta-analysis of data submitted to the Food and Drug Administration.” PLoS Medicine, 5(2), e45. Read the paper ··· This is the source of the “bigger effect than antidepressants” comparison. It’s worth knowing that this paper itself sits at the center of an ongoing controversy.
-
Orben, A., & Przybylski, A. K. (2019). “Screens, Teens, and Psychological Well-Being: Evidence From Three Time-Use-Diary Studies.” Psychological Science, 30(5). Read the paper ··· This is the flagship study on the opposing side. Read it alongside the paper above and you arrive at “we still don’t really know” — and I think that’s the honest place to land.
-
Ministry of Science and ICT & National Information Society Agency (NIA), “2025 Survey on Smartphone Overdependence” (published March 26, 2026) Read the press release ··· The raw data behind the trend where overall numbers improve but adolescents alone get worse.
📝 Glossary
Footnotes
-
PNAS Nexus: A sister journal of the Proceedings of the National Academy of Sciences (PNAS). Launched in 2022 as an open-access journal, meaning you can read this paper’s full text for free. ↩
-
Effect size (Cohen’s d): A standardized measure of how large a difference is before versus after an intervention. By convention, 0.2 is considered small, 0.5 medium, and 0.8 large — though these thresholds are themselves somewhat arbitrary, and the same 0.5 can mean very different things depending on the field. ↩
-
gradCPT (Gradual-onset Continuous Performance Task): A 5-minute task where you press the spacebar when a city photo appears and withhold when a mountain photo appears. Since cities appear 90% of the time and mountains only 10%, pressing becomes habitual — and a moment’s lapse in attention immediately produces an error. It’s widely used to measure attention objectively, without relying on self-report. ↩
-
Meta-analysis: A method that statistically combines multiple studies on the same topic. It’s generally considered more reliable than any single study, but the conclusion can shift depending on which studies are included or excluded. ↩
-
Mediation analysis: A statistical method that goes beyond “A improved B” to estimate a pathway like “A changed C, and that C improved B.” However, mediation analysis cannot prove causation — and the paper itself explicitly notes this limitation. ↩
-
ITT / TOT: ITT (intention-to-treat) is an analysis that includes everyone, even those who didn’t follow instructions, while TOT (treatment-on-treated) analyzes only those who actually complied. ITT is more conservative and realistic — and it’s the standard used for this paper’s main results. ↩
-
At-risk group for smartphone overdependence: A combined category of “high-risk” and “potential-risk” groups as measured by a standardized scale. Keep in mind this is a survey-instrument threshold, not a clinical diagnosis. ↩

Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?