Joy34
10-01 10:31

Two key compute bottlenecks worth keeping in mind, one architectural and the other economic.

Architectural: the memory wall. Compute capacity in FLOPS has gone up massively, but the speed at which data can be fetched from memory like HBM or GDDR has grown much more slowly. Around 55% of AI accelerator bill of materials is memory these days, because more and more has to be added just to keep up with the gap in FLOPS.

Economic: tokens per watt. GPU power consumption keeps rising with each new chip generation. That creates a strong incentive for hyperscalers to raise token throughput using specialized CPUs built around heterogeneous computing. Under that setup, the traditional CPU focuses strictly on application logic and orchestration, while dedicated CPUs handle data movement, security, and specialized math. This is also why $Advanced Micro Devices(AMD)$  is hot in this space with their unified CPU-GPU boxes like the Helios.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment