Every conversation about the AI economy eventually arrives at the same place, which is where all the money is going to end up. Today the attention flows to the model makers and the hardware sellers, with frontier labs commanding extraordinary valuations and the companies supplying the silicon minting some of the largest fortunes in market history, but underneath all of it sits the unglamorous business of turning chips and electricity into usable compute, and I think that layer is about to capture a share of the economics that very few people are pricing in today.
This is a piece about where value accumulates in the AI compute stack over the next several years, and about a specific configuration of supply and demand that hands a small number of operators an extraordinary position. That layer is the cloud sector, the businesses that buy the chips, build the data centers, and rent out the compute. The dynamics I am going to describe lift the whole layer, but the real opportunity sits with the neoclouds, the upstart independents this piece will keep returning to, because they start from a far smaller base than the hyperscalers and therefore have the most to gain as the ground shifts beneath all of them.
I want to be clear at the outset about what kind of argument this is. It is a structural call with a long horizon, not a trade. The dynamics I am going to describe will take years to play out, they will not move in a straight line, and they will be brutal for most of the companies caught up in them. A golden age for a sector is rarely a golden age for every company in it, and this one will be no exception.
Chapter 1: Inception
Every story needs a protagonist, and the protagonist of this one is NVIDIA. That may sound strange for the most valuable company on earth, but the golden age Iâm going to describe grows out of a problem NVIDIA didnât choose and canât solve on its own. The neoclouds are the answer to a question NVIDIA was forced to ask, and that question is what happens when your biggest customers decide theyâd rather not be your customers.
For years the largest buyers of NVIDIA chips have been the hyperscalers, and for years those same hyperscalers have been building silicon of their own. Google has its TPU, now several generations deep and designed alongside Broadcom. Amazon has Trainium and Inferentia, with more than a million Trainium chips deployed and Amazon CEO Andy Jassy describing it as a multibillion-dollar business. Microsoft has Maia and Meta has MTIA. Each of these started as an internal project, a way to shave the enormous cost of running AI workloads by owning more of the stack.
That arrangement is no longer confined to their own workloads. Google now rents TPU capacity to outside AI companies, and Amazon is doing the same with Trainium. The moment a chip built to cut your own costs is offered to someone else, it stops being an efficiency and becomes a product sold directly against the company that used to supply every accelerator in the building.
Anthropic is the clearest illustration of how far this has gone. It committed to up to a million of Googleâs TPUs in the largest deal in Google Cloudâs history, and it runs a supercluster of roughly half a million Trainium chips under a separate arrangement with Amazon. To be fair, Anthropic isnât purchasing custom silicon in either case. Theyâre leasing cloud capacity that happens to run on somebody elseâs chips. But for NVIDIA that distinction changes nothing, because every token Anthropic generates on a TPU or a Trainium chip is a token it didnât generate on an NVIDIA GPU.
Then it gets worse, because the labs are no longer content to solely rent someone elseâs custom chips and several have started designing their own. OpenAI unveiled Jalapeño last month, its first in-house accelerator built with Broadcom and aimed at running GPT inference, part of a ten-gigawatt collaboration announced last year. Anthropic is reported to be in talks with Samsung to manufacture a custom inference chip on a two-nanometer process.
Anthropic has been laying the foundations underneath that silicon as well, signing a twenty-year lease with TeraWulf for roughly 401 MW earlier this month. Whatever hardware fills that campus initially, with NVIDIAâs being the most likely option, a lab that controls the site and the power can swap in chips of its own design as they become ready. Every rack of them that goes in is a rack NVIDIA doesnât fill.
Follow this through and the picture is stark. One way or another, consumers reach AI through the frontier labs or through the hyperscalers, and enterprises reach it through those same two channels. If both channels are increasingly framed in custom silicon, NVIDIA is being slowly designed out of both of its end markets. This isnât a next-year problem, itâs a three-to-five-year one, which is why the decisions NVIDIA makes today matter so much. Get them right and the company ends up as the Apple of AI infrastructure. Get them wrong and it ends up as the Cisco of the dot-com era.
The First Response
NVIDIA isnât a passive company, and it has mounted a response on two fronts. The first is open source.
For the past couple of years NVIDIA has pushed open models harder than any other large player, releasing Nemotron, its own family of open-weight language models that any company can download, run on its own infrastructure, and fine-tune on its own data, alongside the tooling built around the wider open-source ecosystem. On the surface this reads as generosity, but the strategic logic runs straight back to the predicament. The frontier labs are strongest when enterprises have no choice but to rent intelligence directly from them. An open model that a company can run itself, tune to its own data, and own outright is the one thing that genuinely loosens that grip. If open source becomes competitive with the frontier, the labs lose their lock on the enterprise, and the door that was closing on NVIDIA swings back open.
So why would enterprises want that alternative badly enough to matter? Alex Karp has been making the argument for them. In a recent CNBC appearance the Palantir CEO went after Anthropic directly, characterizing how the frontier labs treat the companies building on top of them as predatory, a matter of absorbing your own ecosystem once youâve learned which parts of it are worth absorbing.
The criticism has teeth, and the historical parallel is what gives it force. In the early days of the personal computer, Microsoft sat underneath a thriving market of independent software built on its operating system. Once a category proved valuable, Microsoft had a habit of building its own version and folding it into the platform, which is what happened to spreadsheets, to word processors, and eventually to browsers. The application that pioneered a category kept discovering that the platform beneath it had quietly become its competitor.
Anthropic is running a recognizable version of that playbook. It worked with the coding startups building on its models and then shipped coding tools of its own. It collaborated closely with Figma, whose integration converted Claude-generated code into editable designs, and then launched Claude Design in April, a prompt-to-prototype tool aimed squarely at Figmaâs territory. Anthropicâs chief product officer had resigned from Figmaâs board days earlier, and Figmaâs stock fell on the launch.
Two things are true here at once. Karpâs critique is largely fair, and Karp is also selling something, because Palantirâs ontology product is positioned as the answer to the problem heâs describing. Ontology is an application layer sitting between an enterpriseâs data and whichever model it happens to be running, and because the model never touches that data directly, it cannot cache it, learn from it, or reconstruct what makes the business valuable in the first place. The second consequence matters more here, since a layer built that way makes the model underneath it swappable, and Karp is explicit that this includes open source models running on any cloud provider the customer chooses.
NVIDIA has been happy to ride the same current, partnering with Palantir to push Nemotron into government and giving the open model an institutional channel it wouldnât otherwise have reached. The push has taken policy form as well. This week a coalition of twenty-five companies published an open letter titled Open Weights and American AI Leadership, arguing that Americaâs position in AI should be judged by the strength of its open ecosystem, and the letter reads like this chapterâs argument translated into advocacy. NVIDIA and Palantir sign alongside Meta, Microsoft, Mistral and Hugging Face, and Jensen Huang chose the letter as the subject of his first ever post on X, endorsing it personally while adding the diplomatic note that the world needs both closed and open models. Conspicuously absent from the signatures are the frontier labs themselves, the very companies whose grip on the AI market stands to be threatened by open models.
Whether Nemotron itself ends up mattering is an open question, and it doesnât much change the argument. NVIDIA can keep pouring resources into this specific family or let it fade, and the alignment stays unmistakable either way. NVIDIA wants open source to win a real place in the enterprise, because the alternative is an enterprise served entirely by black-box models running on silicon that increasingly isnât theirs.
The Second Response
Open models still have to run somewhere, and this is where the second front opens and where the neoclouds enter the story. NVIDIAâs interest isnât simply that open source succeeds. Itâs that the infrastructure running it belongs to someone other than a hyperscaler, since the hyperscalers sit among the very companies building the competing silicon. What NVIDIA needs is a home for all that compute that is loyal to NVIDIA by construction, which means an independent cloud layer.
It has spent the past two years building one, and the pattern is difficult to read as anything but deliberate. NVIDIA has taken direct equity stakes in the larger independents, Nebius and CoreWeave among them, tied to priority access on new GPU generations. It structured a deeper arrangement with IREN, taking warrants against GPU deliveries in tranches, which aligns both companiesâ fortunes over the full deployment. At the smaller end it has extended revenue-share financing to a pair of small-cap clouds, where NVIDIAâs involvement is closer to the thing keeping the operation solvent than to a conventional investment. The support runs from the largest independents all the way down to the smallest, which tells you NVIDIA is casting a wide net here, backing the neocloud sector as a whole with very little interest in picking which of them comes out on top.
NVIDIAâs own management describes the motive plainly enough. Asked on the Q4 call about the string of strategic investments the company had been making, Jensen Huang said they were âfocused very squarely, strategically on expanding and deepening our ecosystem reach.â Three months later NVIDIA made the point more concretely by splitting its data center reporting in two. One segment covers the hyperscalers, and the other covers everything else, with the independent AI clouds sitting squarely inside it alongside enterprise, industrial and sovereign buyers. Jensen described that second segment as hundreds of companies today and hundreds of thousands in time, growing at what he called an incredible pace, and pointed out that very few companies have any real exposure to it. Carving out a separate line in the accounts for that customer base tells you where NVIDIA expects its growth to come from.
There is a fair counter-argument to all of this, and it deserves to be put properly. NVIDIA is sitting on an enormous pile of cash and facing a genuine shortage of customers capable of deploying GPUs at the scale it produces them, so investing in the people who buy your product, and helping them buy more of it faster, is about the most ordinary thing a cash-rich supplier can do. Seen that way there is no grand strategy at work, just a chip company financing its own sales channel, and what I have been calling a program is a run of sensible commercial decisions that happen to look coordinated from a distance.
The details are what pull me away from that reading. The support is spread deliberately across many small operators when pure sales optimization would concentrate it in the two or three capable of absorbing the most product fastest. The timing tracks the custom silicon programs far more closely than it tracks demand. And NVIDIA has gone as far as becoming a customer of these clouds, leasing capacity from Lambda and now, in a far larger arrangement, from IREN, to run its own internal workloads and research. NVIDIA could have taken that compute from any hyperscaler and paid whatever was asked without noticing the cost, yet it chose the independents instead. Handing your own AI workloads to a neocloud is about the strongest endorsement available to a company in NVIDIAâs position, and any enterprise weighing whether these operators can be trusted at scale should take note of it.
Whatever the motive, the effect is the same, and the effect is what carries into everything that follows. An entire layer of the compute stack has been capitalized, credentialed, and handed priority access to the best hardware in the world by the company that makes it, and that is the seed. The rest of this piece is about what grows from it, and why the ground it was planted in is about to become far more fertile than most people expect.



