An on-device phone agent may interpret a request, retain context, choose a tool, wait for an app, inspect the result and decide what comes next. That loop repeatedly moves model data, runs transformer operations and coordinates software within tight memory, battery and thermal limits.
That makes on-device agentic AI a system-design problem, not a TOPS contest. Qualcomm’s next-generation Hexagon NPU is a useful example because its announced changes target transformer compute, shared memory, sparse model routing, numerical precision and CPU cooperation. Those features do not establish whether a future phone or agent will be good.
For a product-level definition, read what an AI agent phone is—and what the label does not promise. For device comparisons, see the 2026 global AI phone roundup.
The short answer: an agent needs a balanced pipeline
A local agent needs memory for models and working data, bandwidth to feed its processors, efficient transformer compute, CPU control, fast storage and software that connects the model to permitted app actions. Cooling and power management determine how long it can sustain that work.
| Hardware or software layer | What it contributes | What to verify |
|---|---|---|
| NPU | Accelerates repeated neural-network and transformer operations efficiently | Supported models, precisions and measured latency—not TOPS alone |
| On-chip shared memory | Keeps frequently reused working data close to compute | Absolute capacity, bandwidth and the comparison baseline |
| System RAM and storage | Hold model weights, context caches and experts not resident on the NPU | Available memory after the OS, storage speed and model footprint |
| CPU | Schedules steps, routes work, handles app logic and prepares data | End-to-end task latency and efficiency under mixed workloads |
| OS, runtime and apps | Expose models, tools, permissions, confirmations and fallbacks | Which actions actually work on the shipping device |
| Thermal and battery system | Sustains repeated inference instead of a short demonstration | Long-session performance, heat and energy use |
1. Transformer acceleration reduces the expensive repetition
Most language and multimodal models use transformers. Inference repeatedly applies matrix multiplication, normalization and element-wise functions while processing input and generating tokens. CPUs can perform them, but specialized accelerators can execute common patterns with greater parallelism and energy efficiency.
Qualcomm says its new Hexagon design adds an Element Accelerator aimed at transformer workloads. It works with the NPU’s scalar, vector and matrix extensions rather than replacing them. Think of the NPU as a workshop: matrix units handle large, regular batches of arithmetic, while other units take operations with different shapes and control needs. Adding a dedicated station can reduce handoffs or inefficient use of the wrong machinery.
This is Qualcomm’s architectural claim, not an independent benchmark. It does not show a particular model’s task time, energy use or sustained retail-phone performance.
2. Shared memory matters because data movement is work
An NPU cannot calculate with data it cannot reach. Model weights are learned values, while activations are temporary outputs created as an input passes through the network. Intermediate tensors and model state also have to survive long enough for later operations. Language models may additionally maintain a context cache representing earlier tokens. All of this competes for limited memory.
Qualcomm says the next-generation Hexagon NPU has a 50% larger shared-memory subsystem. Its stated goal is to keep more model state, activations and intermediate tensors on-chip, reducing trips to external DDR memory. A practical analogy is a cook with a larger counter: ingredients used repeatedly can remain within reach instead of being fetched from a pantry for every step. The counter does not enlarge the pantry, but it can reduce waiting and energy spent moving things.
The percentage needs context. Qualcomm’s article does not state the absolute shared-memory capacity in this announcement, and 50% larger does not mean the whole phone has 50% more RAM. It also does not mean an entire large model resides in that NPU memory. Buyers and developers still need total RAM, available RAM, memory bandwidth, storage behavior and workload measurements.
3. Mixture-of-Experts changes how much of a model runs at once
A dense model uses a broad set of its parameters for each inference step. A Mixture-of-Experts (MoE) model instead contains specialized expert blocks and routes each token to a subset. Qualcomm gives an example in which a 30-billion-parameter MoE keeps tens of billions of parameters available but activates about 3 billion routed parameters for each token-generation step on the NPU.
Picture a hospital with 30 departments but only the relevant specialists joining one consultation. The hospital’s full expertise still occupies a building and requires administration; only a fraction of it works on that case. Likewise, “about 3B active” can reduce compute and bandwidth per token, but the rest of the model still has to live somewhere—often across storage and memory—and routing or loading experts has overhead.
Qualcomm describes flash-to-memory expert management and caching as part of the approach. That could make larger model classes practical on a phone, but a 30B MoE should not be presented as equivalent to running a 30B dense model on every token. The example also does not establish model quality, token rate, RAM requirement or battery life for a shipping handset.
4. Precision is a quality, memory and speed decision
Precision describes how many bits represent model values during inference. Lower-bit formats usually reduce weight size, memory traffic and arithmetic cost, but aggressive quantization can hurt accuracy unless the model and runtime handle it well. There is no universally best format.
Qualcomm lists support for INT2, INT4, INT8, FP8 and FP16. That range lets developers assign more compact formats to tolerant parts of a workload and retain higher precision where quality needs it. It is closer to choosing the right image format for each asset than turning one global “fast mode” switch.
The company also claims up to 50% higher prefill performance for INT4 models. Prefill is the stage in which the model processes the prompt and existing context before generating the response; decode is the subsequent token-by-token generation stage. Faster prefill may shorten the pause before output begins, especially with long input. It does not automatically mean the entire agent task is 50% faster. The published claim is vendor-reported, uses “up to,” and does not supply the model, prompt length, power level or comparison platform in the article. Decode speed, tool latency and app handoffs can still dominate.
5. The CPU conducts the agent loop
The NPU is the specialist for model inference, not the operating system’s replacement. An agent still needs control flow: decide which model or tool to invoke, prepare inputs, schedule work, track task state, handle unsupported operations and pass results between apps and processors.
Qualcomm describes the CPU as complementary, orchestrating multi-step workloads and keeping data ready for the NPU. A useful analogy is an orchestra: making the percussion section faster does not decide what the ensemble plays next. End-to-end responsiveness depends on the conductor, the score and the handoffs as well as the fastest section.
Chip capability and product capability therefore differ. A phone maker must integrate the runtime, OS services, app interfaces and permissions. Developers must optimize and test for supported backends. CPU fallback or a missing app interface may outweigh a theoretical NPU advantage.
What better AI hardware does not prove
- Reliability: Faster inference does not show that a model understands a request, selects the right tool or recovers from an error.
- Permission to act: Hardware cannot grant access to messages, payments, files or settings. The OS, app and user must authorize those actions.
- App support: Cross-app work requires supported APIs, protocols or integrations. A capable NPU cannot operate an app that exposes no usable action.
- Safe autonomy: Sensitive actions still need scope limits, transparent previews, confirmation, logs, cancellation and fallback behavior.
- Privacy by default: Local processing can reduce the data sent to a server, but telemetry, backups, account sync and cloud fallbacks may still move data off-device.
- Sustained performance: A short demo does not establish heat, battery drain or speed during a long agent session.
Qualcomm separately argues that moving suitable inference from cloud to edge can improve cost, privacy, performance and personalization. Those are possible design advantages, not guaranteed outcomes. Tasks needing larger models, fresh information or heavy computation may remain in the cloud; practical agents may be hybrid.
Buyer checklist: judge the phone, not the silicon slogan
- Which named agent tasks work today, in your region and language?
- Which steps run fully on-device, and when does the phone use the cloud?
- What happens offline or when an app, account or service is unavailable?
- Which apps expose supported actions, and which require manual handoff?
- Can you review, approve, cancel and audit consequential actions?
- How much RAM and storage are available after the operating system and normal apps?
- Are latency, battery and thermal results measured on a retail device over sustained use?
- How long will the vendor update models, runtimes, security and integrations?
Developer checklist: test the whole action loop
- Profile prefill, decode, tool calls and app handoffs separately.
- Measure peak and sustained memory, bandwidth, energy and temperature.
- Validate quality at each precision instead of assuming lower bits are acceptable.
- Test expert loading and cache misses for MoE models, not only warm runs.
- Record which operations execute on the NPU and which fall back elsewhere.
- Design explicit permission scopes, confirmations, timeouts and recovery paths.
- Benchmark complete user tasks on shipping hardware and software.
FAQ
Is an NPU required for an on-device AI agent?
Not strictly: CPUs and GPUs can run AI models. An NPU is designed to accelerate neural-network inference efficiently, which can matter for battery-powered, repeated workloads. The final experience still depends on memory, software and app access.
Does a 30B MoE mean the phone computes all 30 billion parameters?
No. In Qualcomm’s example, the model has roughly 30 billion parameters available but routes about 3 billion for each token step. The complete model still needs storage and memory management, and routing only a subset does not make it identical to a 3B model.
Does 50% larger shared memory mean 50% more phone RAM?
No. Qualcomm is describing the NPU’s shared-memory subsystem, not total system RAM. The announcement provides a relative increase but not the absolute capacity in that article.
Will faster INT4 prefill make an agent 50% faster?
Not necessarily. Qualcomm claims up to 50% higher performance for one inference phase under unspecified conditions. Token generation, CPU orchestration, network calls, app response time and confirmation steps also affect task completion.
Disclosure and sources
This article analyzes architecture claims published by Qualcomm. The cited performance figures are Qualcomm claims, not independent Meydo tests. Meydo has not used these announcements to assert the capabilities of any Meydo product.
