The massive image: Nvidia lifted the embargo on its Vera CPU deep dive this week, and the disclosure quantities to a direct problem to twenty years of x86 datacenter design philosophy. That is essentially the most element the corporate has shared on the chip because it first appeared on the Rubin roadmap, and it confirms one thing I’ve suspected for some time: Nvidia just isn’t treating the CPU as an connect story anymore. It’s treating it as a battleground with numerous potential {dollars} at play.
Vera is constructed across the Olympus core, the primary custom CPU core Nvidia has ever delivered to the datacenter and the primary customized core the corporate has designed wherever for the reason that Denver and Carmel efforts of the Tegra period practically a decade in the past.
Ryan Shrout is a longtime expertise analyst and business veteran who has spent over twenty years overlaying PC {hardware}, graphics, and semiconductors. He beforehand led technical advertising and marketing at Intel and was the founding editor of PC Perspective. He’s presently President and GM at Signal65. You’ll be able to comply with him on X @ryanshrout.
Grace used licensed off-the-shelf Arm Neoverse V2 cores. Olympus is an Nvidia design from the bottom up, a large, high-IPC core with a 10-wide decode entrance finish that reorders aggressively and prefetches based mostly on patterns like graph constructions in reminiscence.
88 of these cores sit on a monolithic compute die, working 176 threads via a partitioned scheme Nvidia calls Spatial Multithreading, a deliberate departure from the opportunistic useful resource sharing of conventional SMT.
That monolithic alternative issues. Nvidia nonetheless makes use of chiplets for the reminiscence controllers and I/O, however the compute die is one piece of silicon related by a second-generation scalable coherency cloth. The corporate measures bisection bandwidth throughout the die at roughly 3.4 terabytes per second.
The reminiscence subsystem is LPDDR5X hardened for the datacenter with ECC and full telemetry, delivering as much as 1.2 TB/s of bandwidth, roughly 3x the reminiscence bandwidth per core and about 5x the bandwidth per watt of typical DDR-based server designs (all based mostly on Nvidia claims).
The headline claims stack up as roughly 2x sooner efficiency from the Olympus core, 3x the core-to-core bandwidth of chiplet-based competitors, and 40% decrease reminiscence latency below load via the LPDDR5X subsystem.
Vera ships in two varieties: a dense liquid-cooled rack packing 256 CPUs and greater than 22,000 cores, and a standard air-cooled 2U with two sockets. Dell has dedicated to a number of PowerEdge programs constructed on it. Nvidia sizes the chance as a $200 billion growth of the CPU market, which explains numerous the latest market dynamics.
The argument behind the structure
In 2014, a high Xeon carried 14 to 18 cores. At the moment an Epyc Turin half carries 128. Core counts grew roughly 9x over that stretch as a result of cloud economics rewarded rentable vCPUs, whereas per-core efficiency solely about doubled. Chiplets saved prices down however taxed reminiscence bandwidth, information motion, and latency alongside the way in which.
Nvidia argues that agentic AI breaks this commerce. An agent causes on the GPU, then drops to the CPU for software calls, SQL queries, API work, and scripting, then goes again to the GPU, generally lots of of occasions per activity. That loop is sequential.
You can’t throw extra cores at a sequential loop and make it shorter. Solely a sooner core, fed with information sooner, compresses it. The loaded-latency information Nvidia confirmed makes the purpose visually, with chiplet designs hitting a saturation wall simply shy of 400 GB/s of reminiscence site visitors whereas Vera retains scaling.
The corporate has landed on “max single-threaded CPU at scale” because the class title. It’s a mouthful, and I’ll get to that.
The proof factors
The shopper information is early however notable. Perplexity ran coding sandboxes on Vera and accomplished jobs 1.5x sooner than the manufacturing Xeon fleet it runs at the moment, with concurrent sandbox startup 1.9x sooner.
The New York Inventory Trade, which processes 1.1 trillion data a day, examined Vera with the Redpanda streaming engine on HPE programs and measured 6x decrease p99 latency versus Epyc Turin, and is now evaluating it as a alternative. Los Alamos Nationwide Laboratory noticed 7x on an agentic workload and 3x on radiation transport and multigrid simulation codes.
These are actual workloads fairly than artificial benchmarks, which I give Nvidia credit score for, and the supporting documentation places names and configurations on most baselines. They remain vendor-supplied and value studying intently.
The SPEC CPU 2026 numbers carry an estimated label from a pre-production reference system, and each comparability lands on Zen 5 Turin or present and older Intel silicon. The Los Alamos runs had been measured in opposition to a Sapphire Rapids based mostly supercomputer that launched in early 2023. Beating delivery elements is the proper first check, however Venice and Diamond Rapids arrive throughout the 12 months, and that’s the battle that may settle this.
The place there’s extra element wanted
I requested the Nvidia staff immediately through the analyst briefing final week what really separates an agentic CPU from a plain, superb datacenter CPU. We’ve had huge, quick processors working back-to-back loops of VMs and containers for years. Is that this genuinely a brand new workload class, or a quick CPU sporting new advertising and marketing?
To their credit score, the staff acknowledged that “agentic CPU” is the fallacious label and would pigeonhole the half. The trustworthy reply is that the basics haven’t modified, however the charges have.
Agent pipelines hydrate and tear down environments continuously fairly than sometimes. Reminiscence strain is steady. Latency below load turns into the entire recreation, as a result of each millisecond the CPU stalls is a millisecond a really costly GPU sits idle. That may be a actual architectural argument, and the choice to spend die space on per-core pace as an alternative of core depend is a real philosophical break from the place x86 roadmaps have been heading.
What the market wants now’s unbiased, rigorous CPU measurement constructed round these agentic pipelines, run throughout present and subsequent technology elements from each vendor. That’s precisely the sort of work we’re trying ahead to diving into at Signal65.
The ecosystem arrived on day one, to no shock
The accomplice roster connected to this launch is unusually deep for a CPU announcement. OpenAI says it is going to deploy Vera at scale starting in Q3, and the early adopter record additionally consists of Anthropic, SpaceX, and Perplexity, with Los Alamos, NERSC, and TACC representing the supercomputing facet.
Dell, HPE, Lenovo, Supermicro, and Bull all have Vera programs coming, backed by the total ODM bench.
The rack-scale platform is ramping simply as visibly. CoreWeave was the primary cloud to convey up and validate Vera Rubin NVL72 and printed the primary measured numbers from dwell {hardware}, a 10x acquire in tokens per second per megawatt over Grace Blackwell NVL72 on DeepSeek-R1.
Google Cloud stood up the primary A5X occasion on Vera Rubin for the reinforcement studying startup Ineffable Intelligence, Azure and OCI have racks working, and a newly expanded Microsoft and Mistral settlement places Vera Rubin on the heart of a multibillion-dollar European buildout.
On the CPU particularly, DeepInfra, which serves practically 5 trillion tokens every week, measured help for 1.6x extra concurrent brokers and a couple of.2x sooner orchestration versus Granite Rapids. Nvidia counts 300 companions and greater than 350 manufacturing unit websites in 30 nations behind the ramp.
What this implies for AMD, Intel, and the hyperscalers
The timing just isn’t refined. This disclosure lands the day earlier than Advancing AI opens in San Francisco, the flagship AMD occasion the place the 256-core Zen 6 Venice technology of Epyc and the Intuition MI450 household are anticipated to headline the keynote from Lisa Su on Thursday. Nvidia simply set the phrases of the datacenter CPU dialog roughly 24 hours earlier than its largest rival takes the stage.
If a sooner CPU returns GPUs to work sooner, the CPU value turns into a rounding error in rack TCO, and the battle shifts from {dollars} per core to tokens per rack. That framing is the one AMD and Intel now need to reply, and it’s a very completely different dialog than the one which produced 128-core roadmaps.
AMD just isn’t conceding the body. It has already argued that rack-level efficiency per watt favors high-core-count Epyc, and the 256-core Venice technology will sharpen that response. Intel has Diamond Rapids coming. The hyperscalers have Graviton, Axion, and Cobalt, all designed across the scale-out economics Vera explicitly rejects.
However Nvidia just isn’t promoting a service provider CPU right into a commodity socket. It’s promoting the CPU because the utilization lever for the most costly property within the AI manufacturing unit. If a sooner CPU returns GPUs to work sooner, the CPU value turns into a rounding error in rack TCO, and the battle shifts from {dollars} per core to tokens per rack.
That framing is the one AMD and Intel now need to reply, and it’s a very completely different dialog than the one which produced 128-core roadmaps.
Nvidia has promised deeper head-to-head benchmark information in opposition to each x86 and Arm competitors within the coming weeks. That information, and the unbiased validation that ought to comply with it, will inform us whether or not Vera resets the datacenter CPU dialog or just carves out a well-defended area of interest contained in the Nvidia rack.
Source link




