SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture. While GPU clusters handle core model training and forward passes, agentic AI workf

2 min
SpaceXAI Deploys NVIDIA Vera CPUs for Gigawatt-Scale Agentic Infrastructure and Starmind Satellite

SpaceXAI has selected NVIDIA's Vera central processing units to handle the CPU-bound orchestration and execution workloads powering its Grok models as its computing infrastructure expands toward gigawatts of capacity. The deployment spans both ground-based data centers and orbital systems, with SpaceXAI planning to base its first-generation Starmind AI satellite on an optimized Vera Rubin NVL72 rack architecture.

While GPU clusters handle core model training and forward passes, agentic AI workflows create significant processing demands on conventional CPUs. Systems that call external tools, execute code in isolated sandboxes, orchestrate multi-step reasoning chains, and process training datasets frequently stall when host CPUs cannot feed GPUs quickly enough.

NVIDIA Vera Architecture and Orbital Integration

Vera Architecture and Memory Subsystem

NVIDIA designed the Vera CPU specifically for agentic AI workloads, reinforcement learning pipelines, and data processing. The processor features 88 custom Olympus cores equipped with Spatial Multithreading, delivering 176 hardware threads with partitioned execution resources.

The memory architecture represents a departure from standard enterprise x86 designs:

  • Memory Bandwidth: Vera utilizes LPDDR5X memory mounted on detachable, field-replaceable SOCAMM modules, delivering up to 1.2 TB/s of bandwidth per socket.
  • Capacity and Efficiency: The architecture supports up to 1.5 TB of memory per socket while cutting power consumption roughly in half compared to traditional DDR5 configurations.
  • On-Die Fabric: A second-generation on-die mesh links all 88 cores with 3.4 TB/s of bisectional bandwidth, eliminating cross-chiplet interconnect latencies.
  • Coherent Interconnect: NVLink-C2C provides up to 1.8 TB/s of coherent bidirectional bandwidth directly between Vera CPUs and Rubin GPUs.

According to NVIDIA benchmarks, this architecture completes agentic and reinforcement learning tasks up to 1.8 times faster than standard x86 server processors and achieves up to 80 percent faster throughput in containerized sandbox environments. A standard liquid-cooled Vera CPU rack integrates up to 256 processors and supports more than 22,500 concurrent sandbox instances.

Starmind Satellite and Orbital Computing

The deployment extends to SpaceXAI's planned orbital infrastructure. The company plans to construct its first-generation Starmind AI satellite using an adapted NVIDIA Vera Rubin NVL72 system.

The NVL72 platform packages 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 data processing units, linked via sixth-generation NVLink switches. Adapting this liquid-cooled rack-scale system for low Earth orbit requires meeting rigorous thermal dissipation, power distribution, and radiation hardening constraints while retaining identical software abstractions to terrestrial clusters.

Production Deployments

SpaceXAI joins Anthropic, OpenAI, and Oracle Cloud Infrastructure as early production customers deploying Vera hardware. As frontier labs scale computing clusters into the multi-gigawatt regime, reducing CPU bottlenecks in agent execution directly impacts overall token output and operational power efficiency.

Sources

Written by

More to read

  • Reasoning Model Distillation in Production: Trajectory Curation, Thinking-Token Formatting, Over-Thinking Mitigation, and Student RL Alignment

    Reasoning Model Distillation in Production: Trajectory Curation, Thinking-Token Formatting, Over-Thinking Mitigation, and Student RL Alignment Distilling frontier reasoning models into compact language models has emerged as one of the most effective strategies for deploying low-latency, cost-efficient inference pipelines. Rather than training small models purely on input-output answer pairs, reasoning distillation transfers the intermediate exploration, backtracking, and verification trajectori

    1 min
  • Alignment and Uniformity on the Hypersphere: How Geometric Losses Govern Contrastive Representation Learning

    Alignment and Uniformity on the Hypersphere: The Geometric Foundations of Contrastive Representation Learning Contrastive representation learning serves as the foundational objective behind modern neural embeddings, powering dense retrieval systems, visual-language models such as CLIP, and metric learning pipelines. While early literature justified contrastive learning through the InfoMax principle (maximizing mutual information between augmented views), theoretical and empirical analyses have

    1 min
  • Valor and Point72 Back General Intuition at B Valuation for Physical AI and Robotics

    New York-based foundation model startup General Intuition is in discussions to secure new funding at a $6 billion pre-money valuation, according to sources familiar with the matter. The financing round includes new backing from Valor Equity Partners, Point72 Ventures, and Seven Seven Six, alongside continued participation from existing investors Khosla Ventures and General Catalyst. The potential valuation represents a steep increase from the company's previous financing round, which raised $32

    1 min