Viyan

Viyan AI

Inside Apple's M5: Integrating the Neural Engine

Apple has integrated its Neural Engine logic directly into the GPU execution units on the M5 chip to handle dynamic transformer workloads.

Apple has integrated the Neural Engine into the M5 GPU architecture, retiring the standalone NPU block used since 2017. While the A-series and earlier M-series chips relied on a discrete silicon block for neural tasks, the M5 moves this capability into the primary GPU. This is not a matter of simply removing the NPU block; the logic of the Neural Engine—specifically the dedicated multiply-accumulate lanes—has been re-architected to function as part of the GPU's execution units. By embedding these lanes directly into the GPU, the chip designers have removed the boundary between general compute and neural processing.

The Shift to Unified Compute

Previous versions of the Neural Engine relied on a fixed data path designed for predictable, high-throughput tasks. These engines were excellent at processing static convolutional neural networks used in image and video tasks. However, transformer-based models require handling the KV cache—a history of previous token states that changes length with every step of a generation sequence. A static engine struggles here because it expects a fixed data structure, while a GPU excels at dynamic scheduling and parallel memory access. The integration means the M5's GPU now contains hardware-level acceleration for these specific matrix operations, using the GPU's own high-bandwidth memory controller to stream token data.

Feature Standalone NPU (A11-M4) Integrated M5 Architecture
Dataflow Dedicated, isolated path Unified GPU execution flow
Scheduling Static, hardware-level Dynamic, runtime-managed
Memory Path Fixed local cache Shared GPU hierarchy
Primary Target Convolutional (Vision/Image) Transformer (LLM/Decoding)

Solving the KV Cache Bottleneck

To understand why this change matters, consider the memory access pattern of an LLM. In an autoregressive model, the processor must load the KV cache for every new token. Because the cache grows during every generation, a static NPU block often faces stalls, as it cannot easily reallocate memory or shift its workload during a single operation. By folding the neural compute units into the GPU, the M5 allows the chip to treat these neural tasks as standard compute kernels. The GPU's memory controllers use dynamic paging, which lets the processor treat the shifting KV cache as standard memory blocks. This prevents the pipeline stalls common in older, rigid architectures because the GPU can re-fetch data based on the runtime size of the cache without requiring a reset or a costly buffer clear.

Consider an application that performs real-time transcription. On the M4, the system might have struggled to keep the cache resident in the limited local buffer of the NPU, forcing it to frequently swap data to the system memory. On the M5, the GPU-integrated paths allow the kernel to access the cache directly in the Unified Memory, utilizing the GPU's large L2 and system-level caches. This reduces latency significantly during the 'decode' phase of token generation, as the system does not need to synchronize between two separate processing units.

What remains unclear is how the driver stack will handle this shift for existing applications. While Apple's documentation indicates developers should continue using the standard Core ML APIs, the hardware-level change implies that the underlying dispatch logic is fundamentally different. Whether legacy models will see the same performance gains as modern transformer-based workloads depends on how well the compiler can translate older kernel definitions into these new GPU-integrated instructions. We will likely learn more as these chips see wider adoption in heavy inference workloads where the memory management becomes the limiting factor.

Sources