Skip to content
Machine Learning Daily, home

Nvidia pivots to custom silicon with Vera datacenter CPU

Nvidia is challenging two decades of x86 dominance with the introduction of its custom Vera CPU, designed specifically to accelerate agentic AI workloads.

DERRICKCOMPUTE & SILICON727 WORDS

Nvidia has officially detailed its Vera CPU, representing a fundamental departure from two decades of x86-centric datacenter design. This architecture centers on the Olympus core, the first custom-designed CPU core Nvidia has brought to the datacenter market since the Tegra-era Denver and Carmel projects.

The Olympus core utilizes a 10-wide decode front end capable of aggressive reordering and pattern-based prefetching, specifically designed to handle complex memory structures. Nvidia has integrated 88 of these cores onto a single monolithic compute die, employing a technique the company describes as Spatial Multithreading to manage 176 concurrent threads.

This design choice prioritizes high-performance single-threaded execution over the high-core-count strategies currently favored by x86 competitors. The compute die connects via a second-generation scalable coherency fabric, achieving a bisection bandwidth of approximately 3.4 terabytes per second. This monolithic approach minimizes the latency penalties often associated with chiplet-based designs that must traverse multiple interconnects to access shared memory resources.

Nvidia has paired this compute architecture with an LPDDR5X memory subsystem hardened for server environments. This configuration provides up to 1.2 TB/s of bandwidth, which the company claims offers roughly three times the per-core bandwidth of traditional DDR-based server designs. By hardening the memory controller for the datacenter, Nvidia aims to eliminate the bottlenecks that typically plague high-performance computing tasks.

The performance profile of the Vera chip emphasizes reduced latency under load, a critical requirement for agentic AI pipelines. According to data provided by Nvidia, the architecture achieves 40% lower memory latency compared to existing chiplet-based alternatives. This reduction is vital for maintaining high throughput in environments where the CPU must frequently switch between tasks.

The company positions this shift as a response to the sequential nature of agentic AI, where the CPU must frequently handle tool calls, API requests, and database queries. By focusing on per-core speed, Nvidia aims to minimize the idle time of GPUs during these sequential operations. This architectural pivot acknowledges that agentic workflows require a different balance of resources than traditional cloud virtualization.

Early performance metrics from third-party adopters highlight these architectural gains in real-world scenarios. Perplexity reported a 1.5x speed improvement in coding sandboxes, while the New York Stock Exchange observed a 6x reduction in p99 latency when testing the Redpanda streaming engine on HPE systems equipped with Vera. Los Alamos National Laboratory also reported significant gains, noting a 7x improvement on agentic workloads compared to older Sapphire Rapids-based supercomputers.

The significance of this development lies in the changing economics of the modern datacenter. As agentic AI workflows increase the frequency of environment hydration and teardown, the bottleneck shifts from raw core count to the latency of data movement. This shift forces a reevaluation of how compute resources are allocated in large-scale AI clusters.

Ryan Shrout, President and GM at Signal65, notes that the industry has long relied on high-core-count designs to maximize vCPU density. Nvidia is now challenging this paradigm by arguing that the total cost of ownership is better served by CPUs that return GPUs to active computation as quickly as possible. This perspective shifts the primary metric from dollars per core to the efficiency of the entire orchestration loop.

This philosophical break suggests that the future of datacenter silicon may prioritize the speed of the orchestration loop over the sheer volume of cores. The success of this approach depends on whether these performance gains hold up against upcoming x86 generations like Venice and Diamond Rapids. If the Vera architecture proves superior in real-world deployments, it could force a significant shift in how hyperscalers design their next generation of server infrastructure.

Market participants are now looking toward independent, rigorous benchmarks to validate these claims across diverse agentic pipelines. The deployment of Vera at scale, beginning in Q3 with partners like OpenAI, Anthropic, and SpaceX, will provide the first large-scale evidence of its efficacy in production environments. These deployments will serve as the ultimate test for the viability of the Olympus core in high-pressure, real-world conditions.

The immediate competitive landscape remains volatile, with the announcement arriving just before major updates from AMD regarding its Epyc roadmap. Future milestones will center on how hyperscalers integrate these chips into existing racks and whether the performance-per-watt metrics translate into sustained efficiency gains for large-scale AI operations. The industry will be watching closely to see if this custom silicon strategy can effectively disrupt the established x86 ecosystem.

REFERENCED

  1. developer.nvidia.com1.2 TB/s of bandwidth
  2. signal65.comRyan Shrout
  3. tomshardware.comVenice and Diamond Rapids

FILED TO COMPUTE & SILICON

MORE IN COMPUTE & SILICON

ALL

THE DISPATCH

Applied machine learning, filed daily.

Model releases, silicon, clinical deployment, and the policy shaping them. No digest padding.