UncoverAlpha

UncoverAlpha

Nvidia and the Hyperscalers High Stakes Poker Game

UncoverAlpha's avatar
UncoverAlpha
Aug 13, 2026
∙ Paid

The relationship between Nvidia and the big hyperscalers Amazon, Google, and Microsoft was once very clear and friendly, where one was a supplier and the other three were buyers. I believe that era is clearly over, as each is trying to commoditize the layer the other serves. The chip design and data center infrastructure markets. This clash is critical for the AI buildout age, as it determines where much of the AI value (margin) ultimately ends up.

You have Nvidia’s Jensen Huang with the strongest balance sheet in semiconductor history and a ~75% gross margin business. On the other side are the three biggest Nvidia customers — Amazon, Microsoft, and Google — who together will spend around $600B in CapEx this year and have all come to the same conclusion: they do not want to keep paying Nvidia’s margin forever.

In the last few years, the relationship was largely defined by Nvidia being the only game in town for AI chips, and most workloads were used for training. But now things are changing, because:

  • we are shifting to an inference-led market,

  • Nvidia has a historic high 75% gross margin, and it is bothering the data center ecosystem because it’s long-term hurting their margin

  • The stakes have just become so much bigger because hundreds of billions, and maybe even soon trillions, per year will be spent on AI chips, and every layer of the AI stack wants to capture as much value as possible.

Let’s dive in.

Training the land grab. Inference the economy.

Most of you know this well, but just to shortly summarize. Training is when you build the model. You take a giant cluster, you run it for months, and you need the best interconnected hardware because the whole cluster works as one. This is where Nvidia’s moat is the strongest: CUDA, NVLink, the networking stack. In training, Nvidia’s share of the merchant market is still estimated to be north of 90%; the only real competitive chip that comes close is Google’s TPU (mostly because of Google’s interconnect and clustering capabilities).

Inference is when you run the model and is different:

  • It is massively parallel but doesn’t need the same tight cluster coherence as training.

  • The workload is known and stable, which is exactly the situation in which a custom ASIC (a chip designed for one task) often beats a general-purpose GPU in cost per token.

  • It is where most of the future volume is going. Microsoft processed over 100 trillion tokens in a quarter in 2025, up 5x year over year. Google cited 22 billion tokens processed per minute across Alphabet’s APIs.

The customers became the competitors

The problem for Nvidia in this inference world is that the market is so big that companies and their biggest customers don’t want to allow a monopolization of one of the biggest markets in the world; it’s just too much money and profit to be made in this layer, or value transfer to go above the stack for them to just leave it alone. The problem is also that the biggest customers, the hyperscalers, have very strong operating cash flows and business to actually have significant ASIC labs in their conglomerates.

Google is the furthest ahead in chip design. Google produces +3 million TPUs annually. In October last year, it signed Anthropic to a deal for access to up to 1M TPUs, bringing well over a gigawatt of capacity online in 2026. In April this year, that was expanded with Broadcom to roughly 3.5 additional gigawatts of next-gen TPU capacity starting 2027. Google is also now selling TPU systems to external customers as hardware. Essentially doing the same thing Nvidia’s core business is. Because it seems Gemini has fallen behind and the AI lab team is in a refocus period, the likelihood of Google giving even more capacity to GCP and with it TPUs to outside customers has only increased.

Amazon told you on multiple recent earnings calls that it sees its semiconductor design business as one of the next pillars of the Amazon empire. And it’s already significant in size, with big growth rates. Amazon’s chips business (Trainium + Graviton) now runs at over $25B in annualized revenue, growing triple digits. Graviton is used by 98% of the top 1,000 EC2 customers. Project Rainier - the Anthropic cluster is scaling from 500,000+ Trainium chips toward 1M. The problem is also that Amazon, as a company, often likes to drive margins to zero during adoption periods, as they have an obsession with focusing on customers before reaching scale and passing a lot of the economic benefits to their customers. Trainium, given how young the project is, is highly competitive with Nvidia on a cost-per-token basis; much of this is due to Amazon pushing for very low margins. There is a reason why Anthropic likes Trainium for inference tasks.

Microsoft was and still is the laggard, but they, too, have figured out that owning the chip design layer will be key for a hyperscaler. The Information reported that Microsoft will unveil Maia 300 as soon as September and is negotiating with TSMC for capacity for more than 300,000 chips for delivery in 2027, with an eventual target of over 1 million units and, per an Azure Maia manager, “gigawatts’ worth” of Maia. For context, Maia 200 (launched in January 2026, TSMC 3nm, 140B+ transistors) has shipped only in the tens of thousands and is housed in two US data centers. Reports mention a targeted cost advantage of 30-40% over Nvidia chips, and Microsoft is actively courting Anthropic as an anchor customer for it.

The fact that Anthropic has so far become the biggest success story from the AI labs' perspective is also not great for Nvidia, as Anthropic has notoriously built most of its chip-usage stack on Google’s TPUs and Amazon’s Trainium (not Nvidia). Yes, even Mythos (Fable) is rumored to be mostly TPUs and served on Trainium, though that’s also changing with some new Nvidia deals.

Even with Nvidia now trying hard to offer its chips to Anthropic in its deal announcements, it’s clear Anthropic doesn’t want to drop off TPUs and Trainiums. Instead, they want as many options as possible and to drive prices down, which is exactly the scenario Nvidia doesn’t want. Matching workloads to the most suitable, cheapest chip.

And I haven’t even mentioned AMD, where the OpenAI deal ( 6GW of MI-series deployment with equity warrants attached ) and the Broadcom/OpenAI 10GW custom program give buyers even more options, not to mention the SRAM players like Cerebras.

Hyperscaler comments that increased tensions

On the Q2 calls, all three hyperscalers and Meta described their CapEx in a way that Nvidia surely doesn’t like. They framed hundreds of billions in future AI spending as an option rather than a must.

They all committed to long-lead-time CapEx (land, shells, power, cooling — assets that last 15-25 years) years in advance, but said they are going to buy the chips, the “short-lived assets,” only months before a data center comes online. If demand is there, you buy them; if not, you don’t buy the chips.

“if the demand environment changes, you just slow down what is, in fact, the largest component”

(Amy Hood, Microsoft)

Amazon, Meta, and Google essentially said the same thing. With Jassy talking in detail about “different capital cycles” for data center land, power vs chips, etc.

“Google will keep investing as long as we see an attractive return on that investment”

(Google’s Anat Ashkenazil).

The fact is that as more time passes, the more mature the ASIC programs from these companies become, and the more independent they can be in that layer as they get more TSMC allocation or even strike deals with Intel or Samsung for fabrication of these chips. All of the hyperscalers are also going direct to memory manufacturers to secure memory capacity, so their long-term intentions are to essentially copy the AWS playbook on the CPU book with Graviton, where Amazon’s own Graviton chip now handles most of the CPU workloads in AWS.

Nvidia doesn’t want to wait for that and wants to front-load as much demand as possible today, when they still have a dominant market position, but they also want to secure their future and commoditize the hyperscalers.

Nvidia’s counter-move

Nvidia is on track to generate roughly $100B+ of annual free cash flow at current run rates. Nvidia is deploying it to manufacture demand outside the hyperscalers’ walls, so that when the hyperscalers exercise their option to buy less, someone else is will buy more.

Save the neoclouds at any cost. In September 2025, Nvidia signed a $6.3B commitment to buy back any unsold CoreWeave capacity through April 2032; a parallel deal with Lambda was worth $1.5B. Then, in January 2026, with CoreWeave, Nvidia went further by injecting $2B in equity. Nvidia became the equity holder, the demand backstop, and the customer of last resort.

Revenue deal. On July 1, 2026, Nvidia formalized the model as the “AI Compute Partnership”: Nvidia acts as a financial backstop, agreeing to rent back unused GPUs at a fixed rate so lenders will finance neocloud buildouts — and in exchange, Nvidia takes a share of the neocloud’s cloud revenue. First neoclouds to join Firmus (up to 170,000 GPUs in Batam, Indonesia) and Sharon AI (40,000 GB300s in Australia). The problem with neoclouds is that their margins and business models are already razor thin, especially because they don’t have massive free cash flows and they essentially need debt to keep the lights going and to finance their buildouts. Even among them, it seems like the ones taking the Nvidia revenue share deal are seen as the weakest, according to this Former Nebius employee:

“Firmus was almost going bankrupt. It was a cloud computing operator, but they faced substantial challenges, both financially and also regarding the executive board. They had problems regarding customers and also investors. This was the background before. Nvidia stepped in. The ecosystem works like this. When one of the major partners on Nvidia is actually struggling, NVIDIA will try to help the partner and find solutions together”

source: AlphaSense

He also shared his thoughts about what he thinks the Neoclouds are thinking about this Nvidia revenue share deal:

“NeoCloud, they hate this model, so I can tell you directly. They don’t like it because Nvidia has a huge leverage on them, and at the same time they paid a far higher price to have, in this case the GPUs. Like I said, you don’t have the high discounts of buying in bulk. You actually have a far lower discount, and at the same time NVIDIA has also control on how you are using the cloud”.

source: AlphaSense

He also explained possible tension with Neoclouds and Nvidia, as, according to him, Nvidia would prefer that Neoclouds serve smaller, more diverse customers rather than just have big offloading deals with the hyperscalers.

And another move they did was on Monday with the $500B financing deal, bringing Wall Street in more deeply. Nvidia announced a partnership with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to raise more than $500B in third-party capital for AI infrastructure, with compute itself used as collateral in SPVs, and Jensen saying that Nvidia retains the option to backstop up to $125B, or 25%, of the deals.

This is in addition to the reported talks to backstop up to $250B of OpenAI’s compute leases at the SB Energy 10GW Ohio campus and to finance ~$350B of OpenAI chip purchases. Given the recent commentary from SpaceX about going “Nvidia exclusive,” I expect a backstop deal or straight-up equity deal is also brewing in the background when it comes to SpaceX and Nvidia. Elon is smart enough to take advantage of a situation if Nvidia feels it needs to help the ecosystem with financing, as he has never lacked ambition for scale and doesn’t fear pulling the risk lever to the max.

Nvidia already disclosed $18.6B invested into private companies and infrastructure funds in Q1 FY27 alone, with the filing admitting some investees “may indirectly purchase or use our products”.

The problem with these kinds of deals is that when we have a compute shortage on paper, things mostly work out, but the moment we don’t have that shortage anymore, a business that pays Nvidia list price for chips and then pays Nvidia a percentage of revenue on top is a business with no margin left. The neocloud business models, in essence, are all a question in terms of whether this is really a sustainable, durable long-term business model or if this is just a byproduct of the current times where there is not enough compute and where some of the biggest compute buyers (hyperscalers) are doing these neocloud deals to buy some short-term capacity needs while they themselves maintain the end-client relationship. Comparing a hyperscaler business model to a neocloud is, in my view, like comparing apples to oranges, as hyperscalers have the data and ecosystem stickiness needed for a durable business. Also, with their scale and now ASIC development, they will control future margins better than someone with only 1 chip supplier.

These moves by Nvidia are important because they do change the short-term nature of the compute ecosystem and could help create a real AI bubble and an overbuild in compute capacity. Nvidia is right when it comes to neoclouds: for them to have durable business models, they need smaller, more diversified clients. And even that raises a question: once the compute shortage eases, and clients can get cheap, available compute from the hyperscalers, why would they even use neoclouds?

So far, the clients are mostly big companies, AI labs, and the hyperscalers themselves, where the risk is even higher when we do get into a capacity overbuild. The first thing hyperscalers will switch off is Neocloud deals, and the big AI labs are already building their own data centers to become independent of both Neocloud and the hyperscalers. The strategy that Nvidia is trying to encourage with Neoclouds about focusing on non-hyperscalers is a good one for sustainability, but the Neoclouds also like headline-reaching deals with the hyperscalers as it helps their stocks in the short-term.

To continue reading this article, you have to subscribe to the paid section. In the paid section, I will cover why Nvidia needs to finance these deals and how Nvidia, Neoclouds, and hyperscalers are positioned for the future.

Why does a potential monopolist need to finance its own demand?

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Rihard Jarc · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture