When you ask an AI feature on your phone to summarize a message or edit a photo, the work happens in one of two places: on the device in your hand, or on servers in a data center. Increasingly, it is a mix of both. Understanding the difference helps you make sense of privacy promises and why some features only work on newer hardware.
Updated September 2026: we added links to Apple’s and Google’s documentation on how their on-device and private cloud systems handle data, and to device requirements.
How each approach works
Cloud AI sends your request over the internet to powerful servers that run large models, then returns the result. Most chatbots, including the full versions of ChatGPT, Claude and Gemini, work this way.
On-device AI runs a smaller model directly on your phone or computer, using processors designed for machine learning, often called neural processing units or NPUs. Features such as offline dictation, photo search and some writing tools can run this way. On Android, Google’s Gemini Nano runs inside a system service that, according to Google, does not store the inputs or outputs of each request.
Head to head
| Factor | Cloud AI | On-device AI |
|---|---|---|
| Capability | Largest, most capable models | Smaller models, narrower tasks |
| Privacy | Data leaves your device | Data can stay on the device |
| Speed | Depends on connection and server load | Fast for small tasks, no network round-trip |
| Works offline | No | Often yes |
| Cost to user | Often subscription or usage-based | Built into hardware you already bought |
| Battery | Light on the device | Can use significant power for heavy tasks |
The privacy question
On-device processing is a genuine privacy benefit, because data that never leaves your device cannot be exposed in transit or stored on someone else’s server. But “AI on your phone” does not automatically mean everything stays local. Many features quietly hand harder requests to the cloud.
Some companies have built systems designed to protect those cloud requests. Apple’s Private Cloud Compute, for example, is designed so that user data is not retained after a response is returned and is “never available to Apple — even to staff with administrative access,” and Apple publishes its server software images so outside researchers can inspect them. Other providers offer their own privacy commitments. Read the specific policy for the feature you are using rather than relying on marketing phrases.
Why hybrid is winning
Most major platforms now use a hybrid model: a small model on the device handles quick, private tasks, and larger requests go to the cloud. This balances speed, privacy and capability, and it lets companies control server costs.
What this means when you buy a device
New on-device AI features often require recent processors with enough memory to hold the model. That is one reason certain features launch only on the newest phones and laptops: Apple Intelligence needs an iPhone 15 Pro or newer, or a Mac or iPad with an M1 chip or later, and many Windows AI features require a Copilot+ PC with a 40+ TOPS NPU. If on-device AI matters to you, check the manufacturer’s list of supported models before assuming an older device will get new features through a software update.
How to check what you are using Look in your device’s settings for AI or intelligence sections, which often explain which features run on the device and which use the cloud, and let you turn cloud processing off.
The best of both
Cloud AI offers the most capable models; on-device AI offers speed, offline use and stronger privacy for everyday tasks. The best experiences combine both, so the useful question is not which is better, but which parts of your data are leaving your device, and under what protections. For more on the broader shift, see our 2026 trends overview.
Sources
- Private Cloud Compute: A new frontier for AI privacy in the cloud, Apple Security Research, June 2024
- Gemini Nano, Android Developers
- How to get Apple Intelligence, Apple Support
- Copilot+ PCs developer guide, Microsoft Learn



