Wednesday, August 26, 2026

Vera Rubin rack

A Vera Rubin rack is
an advanced, rack-scale server system built by NVIDIA to power large-scale agentic AI and complex reasoning workloads.

  • Memory now accounts for a much larger share of the cost of Nvidia’s newest AI systems. According to Morgan Stanley estimates, it could cost roughly $2 million in a $7.8 million Vera Rubin rack, or about a quarter of the total.
  • Vera Rubin (1928–2016) was a pioneering American astronomer whose groundbreaking research on galaxy rotation curves provided the first convincing evidence for the existence of dark matter.

Key Features of Vera Rubin Racks
  • Agentic AI Design: Built to handle multi-step problem solving and long-context workflows.
  • Cable-Free Build: The compute trays slide in with no loose cables, hoses, or fans, connecting through a printed circuit board midplane.
  • Liquid Cooling: Fully liquid-cooled to handle extreme compute densities efficiently.
  • Automated Assembly: The hardware trays are manufactured using automated robotics for rapid deployment.
The Five Rack Types
The Vera Rubin platform combines five purpose-built rack systems to act as a single supercomputer:
  • NVIDIA Vera Rubin NVL72: The primary compute rack integrating Vera CPUs and Rubin GPUs.
  • NVIDIA Vera CPU Rack: Provides dense, liquid-cooled CPU capacity for code execution and tool use.
  • NVIDIA Groq 3 LPX: Optimized for high-throughput decode inference.
  • NVIDIA Vera BlueField-4 STX: Manages storage operations and infrastructure security.
  • NVIDIA Spectrum-6 SPX Ethernet: Handles high-speed, scale-out network connectivity.
  • Generational Comparison: Blackwell Ultra vs. Vera Rubin NVL72

    SpecificationGB300 NVL72 (Blackwell Ultra)VR NVL72 (Vera Rubin)
    GPU Count72 Blackwell Ultra GPUs72 Rubin GPUs
    CPU Count36 Grace CPUs36 Vera CPUs
    CPU Cores72 ARM cores per CPU88 Olympus ARM cores per CPU
    FP4 Inference Performance1.44 ExaFLOPS3.6 ExaFLOPS
    NVFP4 per GPU (Inference)20 PFLOPS50 PFLOPS
    NVFP4 per GPU (Training)10 PFLOPS35 PFLOPS
    GPU Memory TypeHBM3eHBM4
    GPU Memory Bandwidth~8 TB/s~22 TB/s
    NVLink GenerationNVLink 5NVLink 6
    NVLink Bandwidth (per GPU)1.8 TB/s3.6 TB/s
    Rack-Scale NVLink Bandwidth130 TB/s260 TB/s
    Scale-Out NICConnectX-8 (800 Gb/s)ConnectX-9 (1.6 TB/s)
    CPU-GPU InterconnectNVLink-C2C (900 GB/s)NVLink-C2C (1.8 TB/s)

    No comments:

    Post a Comment