What Is an NPU? The AI Chip in Every New Laptop and Phone, Explained
Quick answer: an NPU (Neural Processing Unit) is a dedicated chip built specifically to run AI tasks — like voice recognition, image generation, and live captions — directly on your device, instead of sending that work to a data center over the internet. It's not a replacement for your CPU or GPU; it's a third specialized processor that handles AI inference far more efficiently, using a fraction of the power a CPU or GPU would need for the same task. As of September 2026, every current flagship chip from Intel, AMD, Qualcomm, and Apple ships with one, and Microsoft requires at least 40 TOPS of NPU performance for a laptop to carry the "Copilot+ PC" label.
Main keyword: NPU explained
What NPU actually stands for, and what it does
NPU stands for Neural Processing Unit. The “neural” part refers to neural networks — the mathematical structure behind most modern AI models. A CPU is built for general-purpose, sequential tasks. A GPU is built for the kind of massively parallel math used in graphics rendering (and, as it turns out, AI training). An NPU is built for one specific job: running an already-trained AI model as fast and efficiently as possible, over and over again, without draining your battery.
That job has a name — inference. Training an AI model (teaching it from scratch) is enormously computation-heavy and almost always happens in a data center on specialized GPU clusters. Inference (using that already-trained model to, say, blur your background on a video call or transcribe your voice) is comparatively light — and it’s the part that NPUs are optimized for.
Why NPUs suddenly appeared in everything
NPUs aren’t new — Apple’s Neural Engine has existed since 2017 in iPhones. What changed is that AI features moved from “nice to have” to central to how operating systems work. Windows’ Copilot+ features, Apple’s on-device Siri intelligence, and Google’s Gemini Nano on Android all depend on running AI models locally rather than in the cloud, for two main reasons:
Latency: a response generated on-device is instant. A response that has to travel to a server and back adds a noticeable delay.
Privacy: for AI features that touch sensitive data — your photos, your voice, your screen activity — running the model locally means that data never has to leave your device.
Running these models on a CPU works, but it’s slow and burns battery fast. An NPU can do the same job using a fraction of the power, which is why manufacturers started building them into every mainstream chip rather than treating them as a niche add-on.
How NPU performance is measured — and why the number can be misleading
NPU performance is usually advertised in TOPS (Trillions of Operations Per Second). As of late 2026, the current generation looks roughly like this:
Chip family
NPU performance
Notes
Qualcomm Snapdragon X2 Elite
Up to 85 TOPS
Highest vendor-published NPU figure as of September 2026
AMD Ryzen AI 400 (XDNA2)
60 TOPS
Same architecture as the previous generation, pushed further
Intel Core Ultra Series 3 (“Panther Lake”)
50 TOPS (NPU alone)
Intel also advertises a “platform total” of roughly 180 TOPS combining CPU+GPU+NPU
Apple M5 Neural Engine
~42 TOPS (unofficial)
Apple stopped publishing an exact TOPS figure starting with recent generations
Here’s the catch: TOPS numbers aren’t directly comparable across vendors. Some companies quote the NPU block alone; others quote a combined “platform total” that adds in GPU and CPU AI performance, which produces a much bigger, more marketing-friendly number. The precision used to calculate TOPS also varies (INT8 vs other formats), and none of that accounts for how well-optimized the software layer is. In practice, a chip with a lower TOPS rating but a more mature software stack can outperform a “faster” chip on paper.
The one number worth remembering: Microsoft set 40 TOPS as the minimum NPU performance required for a Windows laptop to qualify as a “Copilot+ PC” — every current flagship chip from Intel, AMD, Qualcomm, and Apple clears that bar comfortably.
What NPUs are actually used for today
Live captions and real-time translation, processed instantly without an internet connection
Background blur and framing in video calls, run continuously without taxing the CPU
On-device image generation and editing (like Photoshop’s Generative Fill running locally, or Windows’ Image Creator)
Voice assistants that respond without a network round-trip — Siri’s on-device intelligence and similar features on Android and Windows
Smart search across your files and photos based on content, not just filenames
What NPUs are not good at
NPUs are optimized for one repeated pattern: running a fixed, pre-trained model efficiently. They’re not built for training new AI models — that still requires GPU clusters, usually in a data center. They also don’t replace your GPU for gaming or your CPU for general computing; think of an NPU as a specialist that takes a specific, recurring job off the other two chips’ plates.
Frequently asked questions
Do I need an NPU if I don't use AI features?
Not really — if you never use Windows Studio Effects, background blur, or on-device AI tools, the NPU in your chip will mostly sit idle. That said, since NPUs now ship as standard equipment in nearly every current chip, you're rarely paying an extra premium specifically for one.
Is a higher TOPS number always better?
Not necessarily. Because vendors measure and report TOPS differently — some quoting the NPU alone, others quoting a combined platform total — comparing raw TOPS numbers across brands can be misleading. Software optimization and memory bandwidth matter just as much as the raw number.
What's the difference between an NPU and Apple's Neural Engine?
Nothing conceptually — the Neural Engine is Apple's name for its own NPU implementation, built into Apple Silicon chips since the A11 Bionic in 2017. Qualcomm calls its version the Hexagon NPU; the underlying purpose is the same across brands.
This explainer is part of DecodeLayer's Innovation & Future series, covering the hardware and technology trends shaping the devices we review. Figures reflect publicly available vendor data as of September 2026 and may change as new chip generations ship.
Up to 40 hours of battery life with ANC on, customizable EQ in the app, and a price tag under $60. Here's where it delivers — and where the gap to flagship headphones still shows.
To provide the best experience, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent may adversely affect certain features and functions.
Functional
Always active
O armazenamento ou acesso técnico é estritamente necessário para o objetivo legítimo de permitir o uso de um serviço específico explicitamente solicitado pelo assinante ou usuário, ou para o único objetivo de realizar a transmissão de uma comunicação por uma rede de comunicações eletrônicas.
Preferências
O armazenamento ou acesso técnico é necessário para o objetivo legítimo de armazenar preferências que não são solicitadas pelo assinante ou usuário.
Statistics
O armazenamento técnico ou o acesso que é usado exclusivamente com objetivos de estatística.Technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance from your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
O armazenamento ou acesso técnico é necessário, para criar perfis de usuário para enviar publicidade, ou para rastrear o usuário em um site ou em vários sites com objetivos de marketing semelhantes.