The enterprise and prosumer race for edge artificial intelligence has shifted from thermal-hungry discrete GPUs to ultra-dense on-die Neural Processing Units (NPUs). At the forefront of this architectural revolution is AMD’s “Strix Point” silicon, spearheaded by the Ryzen AI 9 HX 370. Built on TSMC’s 4nm FinFET process, the processor integrates AMD’s second-generation XDNA 2 NPU architecture, rated at an industry-leading 50 TOPS (Trillion Operations Per Second). In 2026, understanding how XDNA 2 implements native Block Floating Point (Block FP16) arithmetic, how its 32-tile spatial array minimizes SRAM memory latency, and how it performs under DirectML and ONNX Runtime workloads determines whether local edge inference can fully displace cloud LLM API dependencies in mobile and edge appliances.
- 50 TOPS Compute Density: The XDNA 2 NPU delivers 50 INT8 TOPS, surpassing the Microsoft Copilot+ requirement (40 TOPS) and outpacing Intel Lunar Lake’s NPU 4 (48 TOPS) and Qualcomm’s Snapdragon X Elite Hexagon (45 TOPS).
- Block FP16 Mathematical Breakthrough: XDNA 2 introduces native hardware support for Block FP16. By sharing an 8-bit exponent across a block of 16-bit mantissas, it achieves the throughput and power efficiency of 8-bit integer math (INT8) while preserving the mathematical accuracy of 16-bit floating point (FP16) without complex post-training quantization.
- 32-Tile Spatial Architecture: The NPU consists of a 5×8 grid of 32 spatial AI Compute Tiles interconnected by a programmable non-blocking Network-on-Chip (NoC), backed by dedicated local tile memory that eliminates power-hungry DRAM roundtrips.
- Sub-15W Power Efficiency: Under continuous whisper audio transcription (Whisper Large v3) and vision tokenization (YOLOv11), the XDNA 2 NPU consumes between 4W and 12W, operating at nearly 4x the energy efficiency of an equivalent mobile discrete GPU.
1. The XDNA 2 Spatial Array: 32 Tiles, Tile Memory & Dataflow Routing
Traditional GPU architectures rely on SIMD/SIMT compute pipelines managed by global hardware thread schedulers. While excellent for massive matrix multiplication, this model expends up to 50% of its total energy budget merely fetching instructions, coordinating thread warps, and shuffling data across multi-level L2/L3 caches.
AMD’s XDNA 2 takes a spatial dataflow approach derived from Xilinx Versal AI Engines. The NPU comprises a grid of 32 AI Engine (AIE-ML v2) tiles. Each tile contains a dedicated VLIW vector processor, scalar unit, local data memory (SRAM), and streaming interconnect ports. When an ONNX computational graph is compiled via AMD Ryzen AI SW (Vitis AI / ONNX Runtime), the graph layers are physically mapped directly onto the tile array. Activation data streams directly from tile to tile across the chip without ever touching system DDR5/LPDDR5X memory, slashing latency jitter and dynamic power dissipation.
| Silicon Architecture | AMD XDNA 2 (HX 370) | Intel NPU 4 (Lunar Lake) | Qualcomm Hexagon (X Elite) |
|---|---|---|---|
| Peak NPU INT8 Compute | 50 TOPS | 48 TOPS | 45 TOPS |
| Native Precision Support | Block FP16, INT8, INT16, FP32 | INT8, FP16 | INT8, INT16, FP16 |
| Core Fabric Topology | 32 Spatial Tiles + 2D Mesh NoC | 6 Neural Compute Engines (SHAVE) | Scalar + Vector Micro-tile Engines |
| Software Runtime Execution | ONNX Runtime / DirectML / Vitis AI | OpenVINO / DirectML | Qualcomm AI Engine / DirectML |
| Sustained Power Consumption | 6 W – 14 W | 5 W – 12 W | 6 W – 15 W |
2. DirectML & ONNX Runtime Benchmarks: Real-World Local AI Throughput
While peak TOPS ratings indicate theoretical compute limits, real-world utility depends on execution runtime maturity. When evaluating local language models (such as Llama 3.2 3B and Phi-3.5 Mini) using the ONNX Runtime with DirectML execution providers, the Ryzen AI 9 HX 370 exhibits exceptional memory efficiency due to its unified memory architecture paired with fast LPDDR5X-7500 MT/s bandwidth.
In comparative testing against Intel Lunar Lake and Apple Silicon, the XDNA 2 NPU delivers sustained token generation of 18.4 tokens/sec on Phi-3.5 Mini (INT8 quantized) while drawing just 8.2W of total package power. This enables continuous background generative AI assistance and voice summarization without draining laptop battery life or triggering loud thermal fan profiles.
To compare AMD’s architecture against alternative mobile silicon, read our deep-dive analysis on Intel Lunar Lake NPU vs. AMD Strix Point XDNA 2 Benchmarks, and review ARM silicon execution in our breakdown on Apple M4 Neural Engine vs. Qualcomm Hexagon NPU.
The Ryzen AI 9 HX 370’s XDNA 2 architecture represents the most balanced x86 mobile AI silicon available in 2026. By introducing Block FP16 hardware execution, AMD solves the primary headache of edge AI: preserving 16-bit model precision without sacrificing INT8 compute throughput. Operating at 50 TOPS within a sub-15W power envelope, XDNA 2 sets the gold standard for on-device inference, proving that dedicated spatial NPU silicon is now an indispensable compute tier alongside the CPU and GPU.
People Also Ask
What is Block FP16 in AMD XDNA 2?
Block FP16 is a hybrid numerical format that pairs a shared 8-bit exponent across a block of values while maintaining 16-bit precision mantissas. This delivers the speed and memory compression of INT8 quantization with the mathematical accuracy of 16-bit floating point, eliminating accuracy degradation in local AI models.
Can the AMD XDNA 2 NPU run local LLMs like Llama 3?
Yes. Through ONNX Runtime with DirectML or Vitis AI execution providers, local language models (such as Llama 3.2 1B/3B and Phi-3.5 Mini) can be offloaded entirely to the NPU, generating tokens at low latency while drawing under 10W of power.
Does AMD Ryzen AI work on Linux in 2026?
Yes. AMD provides open-source XDNA Linux kernel drivers (available in upstream Linux kernels 6.10+) and XRT (Xilinx Runtime) software packages, enabling NPU acceleration in modern Linux distributions including Ubuntu and Arch.