Issue #282

Apple Spells Out Its On-Device-vs-Server AI Decision Order

Apple's own docs test on-device models on real tasks first, escalating only when limits are hit.

BusinessApple Spells Out Its On-Device-vs-Server AI Decision Order

Apple’s developer documentation lays out a clear sequence for deciding whether to use the on-device model or the server model. Build the feature first with the on-device model, check its quality with an evaluation tool, and only route it to the server once you’ve determined that its reasoning power or context window1 falls short.

Where this instruction sits in the document matters. It isn’t telling you to decide, once and for all, whether your entire product should run on-device. The vendor is saying, first and foremost, that you should build a single feature, measure it, and let the result decide. In other words, the unit of judgment isn’t the service — it’s the task.


Same Code, Different Limits

In the same document, Apple attaches a five-row comparison table. Both models protect user privacy, but only the on-device model works offline. Usage is unlimited on-device, while the server model caps each user at a daily limit. Multi-step reasoning mode runs only on the server, and the two differ eightfold in how much you can feed in at once — 4K tokens versus 32K tokens.

CategoryOn-Device ModelServer Model (Private Cloud Compute)
Offline operationSupportedNot supported
Usage limitNoneDaily cap per user
Reasoning modeNone3-tier support
Context4K tokens32K tokens

When you create a session, changing a single line — the model object — sends the same prompt to the server instead. Tools and instructions carry over unchanged. So the question a dev team actually has to answer isn’t “should we use on-device?” It narrows down to: does this task fit within 4K tokens, does it require multi-step reasoning, and does it need to keep working on a plane with no connection?

Android’s structure is similar. Google bundles summarization, proofreading, style rewriting, image description, speech recognition, and open-ended prompting under ML Kit GenAI, and all of these features run on a shared Gemini Nano model managed by AICore2. Google lists three advantages in its documentation: input, inference, and output are all processed on-device; the features keep working even with an unstable internet connection; and there’s no per-call server cost.

The fine print is further down the page

Scroll to the bottom of the same Google Doc, and the story changes.

GenAI API inference is only allowed while the app is in the foreground. Even a foreground service doesn’t get you around this — you’ll get a BACKGROUND_USE_BLOCKED error back. Any design that summarizes overnight notifications in the background dies right here, in this one line.

There are also per-app inference quotas. Cram too many requests into a short window and you get ErrorCode.BUSY; exceed the daily long-use limit and you get a separate battery-usage-exceeded error. Running inference on-device doesn’t mean unlimited use — the “bill” just gets converted from dollars into battery and quota.

The device list needs checking too. As of September 10, 2026, the supported devices listed in the document are the Pixel 9–11 lineup, the Galaxy S25/S26 and its foldables, and flagship phones from other manufacturers. Mid-range and budget models aren’t on the list. On top of that, the Gemini Nano version installed varies by device — nano-v2, v3, or v4 — and Google explicitly warns that the same prompt can produce different outputs depending on the version, so you should evaluate beforehand. Language support, including Korean, is also noted as depending on device settings and the downloaded model.

The browser side draws an even sharper line. To use Chrome’s built-in AI, the volume holding your Chrome profile needs at least 22GB of free space, plus either a GPU with more than 4GB of VRAM or 16GB of RAM with a 4-core-or-better CPU. The initial model download requires a non-metered connection, and if free space drops below 10GB after downloading, the model gets deleted. Chrome on Android and iOS isn’t supported yet, and as of Chrome 149, the document lists English, Spanish, Japanese, German, and French as the supported input/output languages. Korean isn’t on that list.

For a Korean service, this is where the math takes a hit. The moment you make on-device processing your default path on the web, you have to start by counting what percentage of your users can even flip the feature on.

Measure After the Plateau, Not the Peak

Performance comparisons demand even stricter conditions. A useful reference here is an empirical paper published in March 2026 and revised in June. The researchers loaded a single 4-bit-quantized Qwen 2.5 1.5B model onto four different devices and fired the same 258-token prompt 20 times in a row, one second apart, logging both speed and thermal state.

The iPhone 16 Pro hit a peak of 40.49 tokens/second on the first two runs, then its thermal state began climbing from the third repetition onward, flattening out at 23.67 tokens/second from the 17th run — a 41.5% drop from peak. The Galaxy S24 Ultra fell more gently, from 12.21 to 10.38 tokens/second, a 15% decline — but that gentle curve only appeared with the screen off, context capped at 2,048 tokens, and the prefill3 chunk reduced to 128 tokens. The paper’s authors explicitly note that these numbers do not reflect the default settings an app developer would ship with.

There’s a clear takeaway for product teams here. For any feature that fires off requests back-to-back, the number that matters isn’t the peak — it’s the speed after the curve flattens out. The figures you see in benchmark videos or launch presentations are almost always drawn from the first few runs.

Extrapolating speed directly from chip specs is equally misleading. In the same study, a board fitted with a dedicated 40-TOPS4 NPU produced only 6.914 tokens/second. That’s because the decode phase — generating the answer one token at a time — is bottlenecked by memory bandwidth, not raw compute. In exchange, this board showed almost no variance across all 20 runs while drawing under 2W, with no observable thermal throttling. At 72 seconds for a 500-token response, it’s unusable for conversational use — but that’s exactly the right profile for overnight summarization jobs where nobody’s sitting there waiting.

Power figures deserve one more note of caution. The same paper discloses that it measured per-token energy differently across platforms, and explicitly warns against comparing the three figures directly. Placing GPU-level measurements, whole-device measurements, and battery-gauge estimates taken with the screen off side by side simply doesn’t produce a valid comparison.

Deciding Not to Use It Is Also a Conclusion

If I sort the conditions I’ve covered so far by task type, the split looks like this.

Nature of the taskProcess on-deviceSend to server
Input lengthShort messages, one screen’s worthLong documents, accumulated conversation
Depth of reasoningClassification, extraction, polishingProblems requiring multi-step reasoning
ConnectivityOffline is a requirementConstant connectivity can be assumed
Timing of executionWhile the user is watching the screenBackground, scheduled execution
FrequencyIntermittent requestsContinuous, high-volume processing

As you move right across this table, the advantage of on-device processing shrinks. If your service has none of the tasks on the left side, the conclusion is simple: don’t adopt on-device processing yet.

Even if you choose to adopt it, the cost doesn’t disappear — it just relocates. Since model versions differ across devices, you have to evaluate output quality version by version; you’ll end up building a separate server-side path anyway for users whose devices aren’t supported; and you’ll need to design what to show users once they hit their quota. Apple went so far as to open up an API specifically so apps can display status when a user is approaching their daily limit.

Any promises about privacy or cost should only cover what you’ve actually verified. Saying a task runs on-device means that specific task’s input never leaves the device — it doesn’t mean every other path in the product is automatically secure too. Samsung letting users choose, in Galaxy AI settings, whether data is processed on-device or in the cloud is itself a design that assumes both paths coexist side by side within the same product.

Oswarld’s Lens

I’m quite bullish on on-device AI. But when I went back through the material to check the basis for that optimism, I found that what’s worth getting excited about and what you can actually build into a product right now aren’t sitting in the same place.

What’s worth getting excited about is the direction. The fact that vendors have put device and server behind the same API, swappable with a single line, reads as a signal that this choice is becoming less of an architectural decision and more of a per-feature setting going forward. And the record of a dedicated NPU holding steady across twenty runs under 2W points toward today’s slowness being a matter of configuration rather than a fundamental limit of the approach.

on deviceWhat you can actually use for a product decision right now, on the other hand, is much narrower. The list of supported devices in the documentation, the foreground constraint, the 22GB requirement, the list of supported languages — these are the numbers that determine today’s activation rate, right now. I think optimism about the technology and the judgment call on whether to put this in this quarter’s roadmap need to be built separately. In fact, while writing this piece, I realized I myself had been holding the two together.

Closing

Whether you adopt on-device AI isn’t decided by how you see the future of the technology. It’s decided by which of your service’s tasks finish within 4K, whether they happen while the user is looking at the screen, and what percentage of your users are on supported devices. And all three of these can be checked right now, using nothing more than vendor documentation and published benchmark data.

That said, the numbers here come from documentation at a specific point in time and from a single round of testing. The list of supported devices and languages changes with every update, and the benchmark used one model and one prompt. The step of re-measuring against your own service’s actual tasks is still ahead of you, no matter what.


💬 What tasks in your service, Reader, might be worth moving on-device? Just tell me the input length and when it runs, and we can work through it together.

📨 If a colleague is weighing whether to adopt an AI feature, at least send them this piece’s task-by-task comparison table.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

Related issues

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

각주

  1. Context window: The total amount of input and output a model can take in and process at once. 4K means roughly 4,000 tokens — feeding in a long document whole requires this number to be large. ↩

  2. AICore: A system service that lets Android run generative models on-device and manages their deployment and updates. It lets apps share a single model instead of each downloading its own. ↩

  3. Prefill and decode: Prefill is the phase where the input prompt is read all at once; decode is the phase where the answer is generated one token at a time. Prefill is mostly bound by compute, decode by memory bandwidth. ↩

  4. TOPS: A chip spec expressing how many trillions of operations it can process per second. Since it represents a theoretical maximum under ideal conditions, it doesn’t directly predict real-world response speed. ↩