Nvidia’s Real Lead Has Moved Beyond the GPU With Vera Rubin


For three years, the dominant story about Nvidia fitted in a single line: the company sold the only graphics cards capable of running AI at scale, and that exclusivity was worth a fortune. Then the cloud giants started designing their own chips, and investors began asking how long the advantage would hold. Since the group published its quarterly results last Wednesday, a different reading has taken hold: the real moat Nvidia has dug no longer lies in the GPU itself, but in everything around it. Here is why.

The GPU is no longer the only playing field

First, a definition. A GPU (graphics processing unit) is a chip originally designed for video games, whose massively parallel architecture turned out to be ideal for training and running neural networks. That component is what sent Nvidia’s valuation up tenfold between the start of 2023 and mid-2025.

Que-es-una-GPU-1-1024x630.jpeg Nvidia's Real Lead Has Moved Beyond the GPU With Vera Rubin

Over the past year, the share price trajectory has been far calmer. The reason is well known: Amazon, Google and other data centre operators now design their own accelerators, and competition on raw silicon is intensifying. The market’s question was therefore legitimate: what is left for Nvidia once the GPU becomes a commodity?

The answer lies in one unglamorous but decisive word: orchestration. Once a data centre reaches gigawatt scale, running the whole thing at peak efficiency becomes an engineering problem far harder than raw compute.

Vera Rubin, or the whole car rather than just the engine

Nvidia is currently rolling out its Vera Rubin architecture. The name covers far more than a processor: it pairs the Rubin GPU with a set of specialised units, including the Vera CPU, an inference accelerator the company designates LPX, and dedicated racks for storage and networking.

The metaphor used by TechCrunch is telling: if the GPU is the engine, these elements are the rest of the car. They do not churn through compute tokens; they make sure everything orbiting the GPU works without waste.

SemiconductorsX-2093625660556157025-01 Nvidia's Real Lead Has Moved Beyond the GPU With Vera Rubin

The Vera CPU illustrates the logic perfectly. Its main job is not to compute but to direct data traffic. “Vera is important because there’s only so much memory that you can put in a single server or any sort of compute platform,” explains Jason Hardy, Nvidia’s vice president of storage technology.

The problem is concrete. Data centre memory capacity has grown in step with compute power, which is precisely what enriched memory makers during the second wave of the infrastructure boom. But having the data is not enough: it still has to reach the GPU at the right moment. Otherwise a processor billed at tens of thousands of dollars spends part of its time waiting.

According to Hardy, the acceleration delivered by the Vera CPU reaches up to three times on these routing operations. “So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking,” he says.

Six chips designed together

NVIDIAAP-2094607194243178610-01 Nvidia's Real Lead Has Moved Beyond the GPU With Vera Rubin

For this generation, Nvidia claims what it calls extreme codesign: six components conceived from the outset to work in concert. They include the Vera CPU, the Rubin GPU, a sixth-generation NVLink switch, a ConnectX-9 SuperNIC, a BlueField-4 data processing unit and a Spectrum-6 Ethernet switch.

In the NVL72 configuration, the platform brings together 72 Rubin GPUs and 36 Vera CPUs. The direct link between CPU and GPU, which Nvidia calls NVLink-C2C, delivers 1.8 terabytes per second of coherent bandwidth. The company advertises up to ten times higher inference throughput per watt than the previous generation, and the ability to train large models with a quarter of the GPU count required on the Blackwell platform.

A word on those terms. Inference is the phase in which an already trained model is used, the one triggered every time you ask an assistant a question. Throughput per watt measures how many responses are produced for a given amount of electricity. It has become the industry’s most scrutinised metric, simply because power is now the main limiting factor for data centres.

OpenAI attacks the same problem from the other end

Nvidia is not alone in identifying this bottleneck. OpenAI developed its own chip, named Jalapeño, with a comparable goal but the opposite method.

“We designed Jalapeño to minimize data movement and communication delays,” the company wrote in a blog post earlier this month. “Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end.”

porquettfin-1983996816128651299-01 Nvidia's Real Lead Has Moved Beyond the GPU With Vera Rubin

Where Nvidia optimises circulation between components, OpenAI tries to eliminate the journey by containing the workload within a single chip. Two philosophies, one shared conviction: performance gains will now come from smarter traffic control, not simply from adding processor cycles.

What it means next

It would be unwise to conclude that Nvidia has definitively won. Moving the contest to the systems layer guarantees nothing: the company will have to face rival chipmakers and cloud operators there just as it did on GPUs. The rules have simply changed. Designing a competing GPU now matters less than making an entire system run at peak efficiency.

For the end user, the effect is indirect but real. Every point of efficiency gained on orchestration lowers the cost of producing an AI-generated response. That cost ultimately determines subscription pricing, how generous free tiers are, and how quickly the most demanding features reach the general public.

One further signal is worth noting: compute is starting to trade as a financial asset, with the CME launching futures on it in August. When a resource becomes tradable on markets, it is usually a sign that it has become a genuine industrial commodity.

In the short term, Nvidia retains a clear lead on this new infrastructure layer. The interesting question is no longer who builds the best GPU, but who can orchestrate tens of thousands of chips without wasting a watt.

Do you think energy efficiency will become the real yardstick for comparing AI platforms? Tell us in the comments and find our hardware analysis on Wanda-techs.com.

read also:

Share this content:

Ingénieur passionné et rédacteur web depuis 2018, j'allie mon expertise technique à ma passion pour l'écriture pour partager astuces, actualités et savoirs pratiques avec la communauté.

Post Comment

Vous avez certainement manqué...