What Is an NPU? The AI Chip in Every New Laptop and Phone, Explained

Quick answer: an NPU (Neural Processing Unit) is a dedicated chip built specifically to run AI tasks — like voice recognition, image generation, and live captions — directly on your device, instead of sending that work to a data center over the internet. It's not a replacement for your CPU or GPU; it's a third specialized processor that handles AI inference far more efficiently, using a fraction of the power a CPU or GPU would need for the same task. As of September 2026, every current flagship chip from Intel, AMD, Qualcomm, and Apple ships with one, and Microsoft requires at least 40 TOPS of NPU performance for a laptop to carry the "Copilot+ PC" label.

Main keyword: NPU explained

What Is an NPU? The AI Chip in Every New Laptop and Phone, Explained

What NPU actually stands for, and what it does

NPU stands for Neural Processing Unit. The “neural” part refers to neural networks — the mathematical structure behind most modern AI models. A CPU is built for general-purpose, sequential tasks. A GPU is built for the kind of massively parallel math used in graphics rendering (and, as it turns out, AI training). An NPU is built for one specific job: running an already-trained AI model as fast and efficiently as possible, over and over again, without draining your battery.

That job has a name — inference. Training an AI model (teaching it from scratch) is enormously computation-heavy and almost always happens in a data center on specialized GPU clusters. Inference (using that already-trained model to, say, blur your background on a video call or transcribe your voice) is comparatively light — and it’s the part that NPUs are optimized for.

Why NPUs suddenly appeared in everything

NPUs aren’t new — Apple’s Neural Engine has existed since 2017 in iPhones. What changed is that AI features moved from “nice to have” to central to how operating systems work. Windows’ Copilot+ features, Apple’s on-device Siri intelligence, and Google’s Gemini Nano on Android all depend on running AI models locally rather than in the cloud, for two main reasons:

  • Latency: a response generated on-device is instant. A response that has to travel to a server and back adds a noticeable delay.
  • Privacy: for AI features that touch sensitive data — your photos, your voice, your screen activity — running the model locally means that data never has to leave your device.

Running these models on a CPU works, but it’s slow and burns battery fast. An NPU can do the same job using a fraction of the power, which is why manufacturers started building them into every mainstream chip rather than treating them as a niche add-on.

How NPU performance is measured — and why the number can be misleading

NPU performance is usually advertised in TOPS (Trillions of Operations Per Second). As of late 2026, the current generation looks roughly like this:

Chip familyNPU performanceNotes
Qualcomm Snapdragon X2 EliteUp to 85 TOPSHighest vendor-published NPU figure as of September 2026
AMD Ryzen AI 400 (XDNA2)60 TOPSSame architecture as the previous generation, pushed further
Intel Core Ultra Series 3 (“Panther Lake”)50 TOPS (NPU alone)Intel also advertises a “platform total” of roughly 180 TOPS combining CPU+GPU+NPU
Apple M5 Neural Engine~42 TOPS (unofficial)Apple stopped publishing an exact TOPS figure starting with recent generations

Here’s the catch: TOPS numbers aren’t directly comparable across vendors. Some companies quote the NPU block alone; others quote a combined “platform total” that adds in GPU and CPU AI performance, which produces a much bigger, more marketing-friendly number. The precision used to calculate TOPS also varies (INT8 vs other formats), and none of that accounts for how well-optimized the software layer is. In practice, a chip with a lower TOPS rating but a more mature software stack can outperform a “faster” chip on paper.

The one number worth remembering: Microsoft set 40 TOPS as the minimum NPU performance required for a Windows laptop to qualify as a “Copilot+ PC” — every current flagship chip from Intel, AMD, Qualcomm, and Apple clears that bar comfortably.

What NPUs are actually used for today

  • Live captions and real-time translation, processed instantly without an internet connection
  • Background blur and framing in video calls, run continuously without taxing the CPU
  • On-device image generation and editing (like Photoshop’s Generative Fill running locally, or Windows’ Image Creator)
  • Voice assistants that respond without a network round-trip — Siri’s on-device intelligence and similar features on Android and Windows
  • Smart search across your files and photos based on content, not just filenames

What NPUs are not good at

NPUs are optimized for one repeated pattern: running a fixed, pre-trained model efficiently. They’re not built for training new AI models — that still requires GPU clusters, usually in a data center. They also don’t replace your GPU for gaming or your CPU for general computing; think of an NPU as a specialist that takes a specific, recurring job off the other two chips’ plates.

Frequently asked questions

Do I need an NPU if I don't use AI features?

Not really — if you never use Windows Studio Effects, background blur, or on-device AI tools, the NPU in your chip will mostly sit idle. That said, since NPUs now ship as standard equipment in nearly every current chip, you're rarely paying an extra premium specifically for one.

Is a higher TOPS number always better?

Not necessarily. Because vendors measure and report TOPS differently — some quoting the NPU alone, others quoting a combined platform total — comparing raw TOPS numbers across brands can be misleading. Software optimization and memory bandwidth matter just as much as the raw number.

What's the difference between an NPU and Apple's Neural Engine?

Nothing conceptually — the Neural Engine is Apple's name for its own NPU implementation, built into Apple Silicon chips since the A11 Bionic in 2017. Qualcomm calls its version the Hexagon NPU; the underlying purpose is the same across brands.

This explainer is part of DecodeLayer's Innovation & Future series, covering the hardware and technology trends shaping the devices we review. Figures reflect publicly available vendor data as of September 2026 and may change as new chip generations ship.

Get the next review