NVIDIA's New Rubin Server Racks Ease Cloud Deployment Challenges Ahead

Deep News07-28 17:22

Cloud service providers have recently faced numerous obstacles in building data centers, but a positive development is now on the horizon. NVIDIA's next-generation chips are expected to present fewer deployment difficulties.

Several cloud executives who purchased NVIDIA AI servers have confirmed that after initial testing, they anticipate the upcoming Vera Rubin server racks will be significantly easier to install than the 2025 Grace Blackwell racks. Last year, NVIDIA's top customers were overwhelmed by a host of unfamiliar new technologies.

These deployment issues were kept confidential until recently brought to light. Cloud providers invest billions in hardware, and every week spent troubleshooting servers delays the launch of high-yield computing resources for AI clients. Delays also impact data center contractors' order fulfillment and reduce their profit margins.

Currently, the Vera Rubin racks are in limited early customer delivery. The significant test of large-scale deployment will occur next year, when customers receive hundreds or thousands of racks in bulk. NVIDIA plans to release a more powerful upgraded version in the second half of next year, but this high-end model may introduce new problems due to substantially higher power consumption and a greater number of required switches and cables for networking.

The base version of the rack currently under testing contains 72 GPUs and 36 CPUs per unit. It is essentially NVIDIA's third-generation mature product, not a brand-new architecture, with hardware configurations highly similar to the Blackwell racks. A high-performance version supporting 144 or 576 GPUs is not yet in production.

A cloud industry executive stated, "Frankly, we don't think deploying Vera Rubin at scale will be too difficult. The mechanical structure and cooling solutions are already in their third generation and are quite mature."


Software and Hardware Maturity Contrasts

NVIDIA claims the Vera Rubin system is highly compatible with existing software for Blackwell servers. However, two cloud executives admitted that companies will still need months to debug software for stable operation of large Vera Rubin clusters. The industry jokingly refers to the deployment process as a "mystical operation."

The purchase cost of a Vera Rubin rack is at least double that of the first-generation Blackwell rack. Nevertheless, NVIDIA data shows a tenfold improvement in efficiency, measured as tokens per second per watt of AI output.

Installation of the Vera Rubin rack is still not easy. One cloud executive noted that many components are brand new, with a total of 1.3 million parts per rack, compared to 1.2 million for the first-generation Blackwell. At the GTC developer conference in March, NVIDIA CEO Jensen Huang showcased a new network switch for interconnecting Rubin chips, describing the hardware as "incredibly difficult to develop."

Huang did not specify if this new network would create deployment obstacles for customers, but similar network technology delayed the large-scale rollout of Blackwell racks last year. An employee at Oracle, who manages large computing clusters including upcoming Rubin servers, said troubleshooting network faults in these clusters is extremely difficult, relying on experience and described as "mystical troubleshooting."


Power and Cooling System Upgrades

The power consumption of a single Rubin rack is 75% higher than a Blackwell rack, generating more heat. Interestingly, the new rack's cooling system itself has lower power consumption than Blackwell's, but its installation is more complex. A long-time NVIDIA analyst suggested that cooling engineering issues delayed the initial mass production of the first Vera Rubin servers by about two months.

An NVIDIA spokesperson stated, "Product planning remains consistent with previously disclosed information, with no changes." Beyond the hardware, successful deployment of Rubin servers also depends on customers building new, dedicated data centers following NVIDIA's recommended layouts, cabling, and management software for efficient rack deployment.

Kevin Cochrane, CMO of GPU cloud service provider Vltr, stated that customers must build entirely new infrastructure to accommodate the Rubin equipment, as retrofitting existing older data centers is too costly and uneconomical.


Ultra High-Performance Version Poses New Challenges

The upcoming Vera Rubin Ultra high-end rack, set for release next year, introduces several new designs that will require significant time for adaptation and debugging. Unlike standard server trays that slide horizontally into the rack, the Ultra model's chip trays are inserted vertically, like books on a shelf. This changes the gravitational load on the entire system.

The Ultra rack can interconnect up to 576 GPUs and has nearly three times the power consumption of the standard version, which could create new deployment issues. Andrew Bell, NVIDIA's senior vice president of hardware, stated earlier this month that the Blackwell architecture essentially rebuilt the entire GPU system from scratch, causing significant customer deployment pain. In contrast, the Vera Rubin changes are more moderate, "massively reducing manufacturing complexity."

Bell provided data showing a 95% first-pass assembly yield for the Rubin chip tray, compared to only 20% for the first-generation Blackwell. However, not everyone shares this optimistic view. Another cloud executive expressed doubt, believing the high-performance Rubin will still face deployment challenges similar to Blackwell.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment