A Vera Rubin rack is an advanced, rack-scale server system built by NVIDIA to power large-scale agentic AI and complex reasoning workloads.
- Memory now accounts for a much larger share of the cost of Nvidia’s newest AI systems. According to Morgan Stanley estimates, it could cost roughly $2 million in a $7.8 million Vera Rubin rack, or about a quarter of the total.
- Vera Rubin (1928–2016) was a pioneering American astronomer whose groundbreaking research on galaxy rotation curves provided the first convincing evidence for the existence of dark matter.
Key Features of Vera Rubin Racks
- Agentic AI Design: Built to handle multi-step problem solving and long-context workflows.
- Cable-Free Build: The compute trays slide in with no loose cables, hoses, or fans, connecting through a printed circuit board midplane.
- Liquid Cooling: Fully liquid-cooled to handle extreme compute densities efficiently.
- Automated Assembly: The hardware trays are manufactured using automated robotics for rapid deployment.
The Five Rack Types
The Vera Rubin platform combines five purpose-built rack systems to act as a single supercomputer:
Generational Comparison: Blackwell Ultra vs. Vera Rubin NVL72
| Specification | GB300 NVL72 (Blackwell Ultra) | VR NVL72 (Vera Rubin) |
|---|---|---|
| GPU Count | 72 Blackwell Ultra GPUs | 72 Rubin GPUs |
| CPU Count | 36 Grace CPUs | 36 Vera CPUs |
| CPU Cores | 72 ARM cores per CPU | 88 Olympus ARM cores per CPU |
| FP4 Inference Performance | 1.44 ExaFLOPS | 3.6 ExaFLOPS |
| NVFP4 per GPU (Inference) | 20 PFLOPS | 50 PFLOPS |
| NVFP4 per GPU (Training) | 10 PFLOPS | 35 PFLOPS |
| GPU Memory Type | HBM3e | HBM4 |
| GPU Memory Bandwidth | ~8 TB/s | ~22 TB/s |
| NVLink Generation | NVLink 5 | NVLink 6 |
| NVLink Bandwidth (per GPU) | 1.8 TB/s | 3.6 TB/s |
| Rack-Scale NVLink Bandwidth | 130 TB/s | 260 TB/s |
| Scale-Out NIC | ConnectX-8 (800 Gb/s) | ConnectX-9 (1.6 TB/s) |
| CPU-GPU Interconnect | NVLink-C2C (900 GB/s) | NVLink-C2C (1.8 TB/s) |

No comments:
Post a Comment