Two key compute bottlenecks worth keeping in mind, one architectural and the other economic.
Architectural: the memory wall. Compute capacity in FLOPS has gone up massively, but the speed at which data can be fetched from memory like HBM or GDDR has grown much more slowly. Around 55% of AI accelerator bill of materials is memory these days, because more and more has to be added just to keep up with the gap in FLOPS.
Economic: tokens per watt. GPU power consumption keeps rising with each new chip generation. That creates a strong incentive for hyperscalers to raise token throughput using specialized CPUs built around heterogeneous computing. Under that setup, the traditional CPU focuses strictly on application logic and orchestration, while dedicated CPUs handle data movement, security, and specialized math. This is also why $Advanced Micro Devices(AMD)$ is hot in this space with their unified CPU-GPU boxes like the Helios.
Comments