Executive Summary
The stake. Roughly $1.8 trillion of private-market value is marked to the bet that frontier labs OpenAI and Anthropic stay ahead — Anthropic at $965B, OpenAI at $852B. Model pricing power and exclusive training data are the two most consequential concentrations—both currently at risk and are examined in detail in this issue.
A future issue of HyperDensity will treat a third layer AI concentration—compute and power layers beneath them.
The reel. In four weeks, eight AI leaders staked public positions on open vs. closed weights. Nvidia, Meta, Google, xAI, Microsoft, and Palantir verbally back open (some of their actions show a more hedged approach). OpenAI and Anthropic raise concerns about open models. While promoting philosophy, each also promotes their book — open weights models are a more viable business proposition for legacy tech companies with existing, alternative revenue than they are for the two emerging frontier labs whose revenue is mostly generated by direct, closed model access.
The Precedent. There is a historical precedent for how open source technology accrues value. Open models in AI present a new construct, however: open weights. This issue explores the implications of open source vs. open weight vs. closed paths, and why value capture via open technology may be different this time.
The macro layer. Apollo’s Torsten Slok: silicon runs at a 41% gross margin, models and applications at a negative 59% operating margin. The most profitable layer depends on the least profitable one continuing to raise capital. Global AI investment is projected above $1 trillion in 2026; the BIS has flagged the buildout as a macro risk. Everyone—including frontier model competitors—needs Anthropic and Open AI to succeed.
The first concentration within frontier labs is model pricing power. The bear case: The UK’s official AI safety institute finds open weights now trail the closed frontier by only four to seven months, at five to fifty times lower per-token cost. If capability converges while price stays this cheap, the $1.8T mark potentially compresses. The bull case: Brad Gerstner via Jensen Huang: enterprises pay ~5x more for closed frontier because the API absorbs the productization work—frontier models are worth the premium.
The second concentration within frontier labs is exclusive training data. A shrinking pool of licensors — Reddit, News Corp, FT, Stack Overflow, Shutterstock — supplies what frontier labs can legally acquire, and premium data is consolidating into exclusive deals between a handful of publishers and a handful of closed labs. The legal reckoning ($1.5B Anthropic settlement, landmark fair-use ruling, publisher suits) makes the scrape-first playbook unrepeatable and favors incumbents. The bull case: Amodei’s push for safety standards, transparency, and licensing regimes mature the industry past its Wild West years, an interpretation that preserves frontier advantages in data and pricing power. The bear case: Anthropic’s regulatory push raises the bar on open rivals catching up. If regulation ultimately favors open weight, however, it could embolden open models that distill—some say steal—closed models data/methods, and, as a result, compress frontier pricing power. Which one wins in Washington over the next 18 months decides which way the mark moves.
Open Season
July 12. CNBC studio. Palantir’s Alex Karp — known for saying in public what other CEOs say in private — tells a live camera that “something has gone completely wrong” with frontier AI. The clip goes viral over the weekend. By Monday, every AI CEO in America is being asked to respond.
July 24. Nvidia. Jensen Huang publishes his first-ever post on X — an industry letter from twenty-five tech companies defending open-weight AI. “The world needs both.” Two names are missing: OpenAI and Anthropic.
Same day. Microsoft. Satya Nadella hosts the letter on Microsoft’s own site and calls open weights “essential.” Given Microsoft’s long association with OpenAI, the signature reads as a hedge as well as a statement of principle.
That weekend. OpenAI is missing from the letter. Twitter notices. By Sunday, OpenAI’s name is added. Sam Altman says nothing. Elon Musk — still suing him for violating OpenAI’s original agreement to remain open source — quote-tweets a fresh Apple lawsuit against OpenAI claiming trade-secret theft and IP misappropriation: “Scam Altman strikes again.” Musk has pledged to open Grok by year-end. Only Grok-1, from 2024, has actually shipped that way.
July 27. Anthropic. Dario Amodei — pressed all week to explain Anthropic’s cautious posture toward open weights — publishes his response. He advocates chip controls, limits on open-source models training from larger models (AI distillation), and safety testing for all frontier models. Those are defensible safety positions. They could also raise compliance costs for open-weight competitors, a trade-off any policy debate should acknowledge. Meanwhile, Sundar Pichai says little. Google’s open Gemma model has been downloaded 900 million times, Gemini remains closed, and Google is on the letter.
August 10. Meta. Mark Zuckerberg publishes The Future Is for Everyone and reframes the whole debate around balance of power. He says Meta will resume releasing open models. The word “resume” is the news: it concedes that Meta had paused its open stance.
Most people watching all this see a set of AI leaders arguing about a philosophy. They are also contesting something more concrete: who controls the substrate of the next economy.
The answer will shape where the next decade of enterprise value concentrates — and, in turn, the power contracts, data centers, software budgets, and balance sheets built around it. It will also define the course of markets, industries, and businesses which are, like it or not, now thrust into the AI story.
No one knows how it resolves. But open versus closed systems have an instructive history which might serve as a guide.
1. The IBM lesson
In December 2000, IBM announced it would invest $1 billion in Linux — the open-source operating system built by volunteers and mocked in Redmond as a communist toy. IBM reassigned 1,500 programmers to Linux development and put another $300 million into enterprise consulting and support.
Wall Street was skeptical. IBM was subsidizing free software that competed with its own AIX operating system. What was the return?
By 2003, Linux-related revenue at IBM was over $2 billion. IBM Global Services — the consulting arm that helped enterprises deploy Linux, integrate it with legacy systems, and manage the underlying hardware — grew from $33 billion in 2000 to $47.4 billion in 2005.
IBM was not the technology leader of the next twenty years — nobody would confuse it with Google, Amazon, or Nvidia. Some readers will see a company that lost its innovation edge. Fair. But Linux helped protect the business through harsh competition. Survival is not the ceiling of ambition. It is the floor.
And IBM is not alone
The playbook appears repeatedly once you know where to look.
Google + Android (2007). Google open-sourced Android for handset makers, then captured default search, the Play storefront, and the advertising rails around both. Google never needed to sell the operating system.
Google + Chromium/Chrome (2008). Google open-sourced Chromium and shipped Chrome on top of it. Chrome’s default search engine fed the advertising business above the browser.
Databricks + Apache Spark (2013). The founders open-sourced Apache Spark — the distributed data-processing engine they had built at Berkeley — then built Databricks as the commercial platform on top. Spark became the default engine for large-scale analytics; Databricks captured the enterprise deployment layer around it and reached a $188 billion valuation in July 2026.
Same pattern each time: a for-profit company supports the open layer it cannot own and captures adjacent layers it can.
But the playbook fails when you pick the wrong layer
In 2007, Sun Microsystems open-sourced Java on the same theory as IBM: support the substrate, capture the adjacent layers. Sun lost not because open source failed but because every adjacent layer it hoped to monetize was already being taken from it — Solaris by Linux, its servers by commodity Intel hardware, Java itself by other platforms (most consequentially Google’s Android). Two years later Sun sold to Oracle for $7.4 billion, a fraction of peak value.
That is the risk every AI leader in the reel is trying to avoid: supporting the substrate, then discovering five years later that there is no durable layer left to monetize.
2. Why AI needs a different version of the trick
IBM gave Linux away under the GPL — a license that permits source access, modifications, and redistribution. It worked because Linux itself would have been difficult to monetize directly, and IBM’s real revenue generators — services, integration know-how, and sales force — lived elsewhere. The substrate and the moat were different objects.
AI is not like that. The trained model may carry proprietary assets that model creators do not want to release, including:
The training-data corpus. User-behavior, licensed, and specialized data can be proprietary; release can expose signal others can exploit.
The post-training recipe. The techniques that turn a raw model into a usable one — reinforcement learning from human feedback (RLHF), constitution-style rulebooks, and safety alignment work — are hard-won and rarely shared.
Legal exposure on inputs. Rights to some training material are contested; release could make disputed inputs easier to identify.
To provide open access while addressing these sensitivities, the AI industry has developed a middle option: open weights.
Open source (strict). Weights, training data, full training code, permissive license. Almost no frontier model qualifies. The Open Source Initiative has ruled Llama non-compliant twice.
Open weight. The trained model is downloadable and runnable. Training data stays closed, and licenses often carry restrictions. Llama, Qwen, DeepSeek, gpt-oss, Nvidia’s Nemotron, and an early version of Grok fit here.
Closed. No weights released; access comes through an API. GPT-5.x (except gpt-oss), Claude, current Grok models, and Gemini fit here.
Open weight is the AI-native version of IBM’s Linux move. It can broaden access to the substrate while keeping training data, post-training methods, and legal exposure inside the vault. Whether and how model providers prove — beyond contracts and attestations — that their data, which includes both training data and their customers’ data, truly is private is a lengthy post for another day in the coming weeks.
Why the eight leaders line up where they do
Nvidia, Meta, Google, xAI, Microsoft, and Palantir have revenue engines outside the model itself, making open weights a more natural fit. OpenAI and Anthropic face a different trade-off: direct model access remains central even as both build applications and enterprise distribution around it. The incentives differ; neither posture is inevitable.
The question is whether OpenAI and Anthropic can build revenue engines beyond the model itself — or keep their models far enough ahead that direct access stays the product. Both paths are legitimate; neither is assured. Anthropic’s May Series H valued it at $965 billion post-money; Reuters put OpenAI’s March valuation at $852 billion. Those prices leave little room for a strategic miss.
On a recent All-In podcast, Chamath Palihapitiya argued that pricing Anthropic and OpenAI as if their model advantage will last a decade is a “mathematical mistake.” David Sacks conceded on the same podcast that both face a real question about how fast their edge gets copied.
That is one thesis, not a settled outcome. Closed labs may sustain a capability lead, turn access into workflow lock-in, or build platform economics beyond APIs. The live question is not whether closed labs are doomed. It is where their durable advantage, if it persists, will reside.
3. And it matters far beyond AI
Apollo’s chief economist Torsten Slok published the clearest version of this argument on August 7. Slok divides the AI value chain into four layers: silicon and equipment (Nvidia, AMD, Micron) runs at a 41% gross margin. Models and applications (Anthropic, OpenAI) runs at a negative 59% operating margin. Compute and energy sit in between.
Slok’s conclusion: “The most profitable part of the AI value chain depends on the least profitable part continuing to grow revenue or raise capital.” Upstream margins may be real, but they are financed by a layer still raising capital rather than generating sufficient end-demand cash flow. Capital can bridge that gap. Not indefinitely.
The scale explains why this is not only an AI story. Goldman Sachs projects global AI investment above $1 trillion in 2026. The five major hyperscalers issued $121 billion in debt in 2025, four times their prior five-year average. The Bank for International Settlements warned in June that their AI investment was outpacing earnings and free cash flow, and flagged that disappointment in returns could turn the capex boom into a protracted investment bust with knock-on effects on financial conditions. Oracle alone carries $130 billion in debt against $260 billion in not-yet-started AI-infrastructure lease commitments and a $300 billion OpenAI agreement.
That makes the substrate contest a macro story. If model economics compress and end-customer ROI arrives too slowly to justify upstream spending, the correction runs through the same balance sheets that carry the broader market. Data-center REITs, power-generation project finance, GPU lenders, and enterprise-SaaS multiples are all exposed to the pace at which that ROI appears. The question is no longer whether AI spending is large. It is whose balance sheet carries the gap between construction and proven demand — and that answer turns partly on who wins the substrate war.
One silver lining under the fight: while Nvidia, Palantir, Google, Meta, Microsoft and xAI all back some version of open weights, their businesses still ride on OpenAI and Anthropic succeeding. Closed labs are the largest single source of GPU demand, cloud consumption, and application partnerships that make the open-weight ecosystem viable at scale. Everyone in the reel needs frontier models to succeed.
With so much at stake, it is important to laser focus on two high stakes concentrations that are currently under fire within frontier models.
4. Concentration #1 — Model Pricing Power
The most consequential concentration in models is not the models themselves. It is the roughly $1.8 trillion of private capital marked to the bet that two of them stay ahead. And that bet, in the end, is a bet on pricing power — on whether OpenAI and Anthropic can keep prices high enough to fund the capex, but low enough that enterprises don’t bother wiring up their own open-weight stack.
Anthropic was valued at $965 billion in May. Reuters put OpenAI’s March valuation at $852 billion. Between them, that is nearly two trillion dollars of private-market value marked to the thesis that the model layer stays a durable moat, not a commodity.
The adoption case is real. Anthropic’s annualized revenue went from $87 million in January 2024 to a company-disclosed $47 billion by May 2026, and to a tracker-estimated $74 billion in July — the steepest enterprise-software ramp anyone in the room has seen.
Sources: Anthropic’s enterprise update, Reuters and its 2025 outlook, Anthropic’s April update, and Series H. The July figure is a third-party tracker estimate, not a company disclosure. An annualized run rate is the latest month’s revenue extrapolated forward — it signals adoption, not recognized annual revenue or cash flow after compute costs.
Update, August 14, 2026. After this piece was drafted, Bloomberg reported that Anthropic's preliminary Q2 2026 revenue exceeded $11.5 billion — up from $4.73 billion in Q1 and $787 million in Q2 2025, a 14-fold year-over-year jump — and that the quarter produced positive adjusted operating income. That is quarterly recognized revenue, not an annualized run rate, and it is the cleaner number: it is what enterprises actually paid Anthropic between April and June. The IPO Anthropic filed for confidentially is now targeted for this fall, with Morgan Stanley, Goldman Sachs, and JPMorgan advising; some investors are floating a $2 trillion valuation. The adoption case just got stronger, and one bear line needs a caveat: Slok's "negative 59% operating margin" on the Models and Applications layer aggregates the whole layer. In Q2, at least one of the two labs cleared the line. Whether that persists at $47B+ run-rate scale — where compute costs, safety spend, and enterprise sales-and-marketing scale together — is the next thing to watch.
The bear case is also real. The UK’s official AI safety institute recently found that the best open-weight models now trail the closed frontier by only four to seven months, down from six to ten months a year earlier. And they are dramatically cheaper: strong open-weight models typically run five to fifty times less per token than closed frontier flagships, with the widest gaps showing up in output-heavy workloads where most bills are decided. If that gap keeps closing on capability while the price gap stays this wide, closed-model access commoditizes, and the concentrated capital that has been priced against a wide moat has to find a narrower one.
The bull rebuttal from Brad Gerstner on All-In: raw token price is the wrong unit. Enterprises pay roughly five times more for the frontier because — as Jensen Huang has argued through Gerstner — a closed-lab API absorbs the training, fine-tuning, safety, and maintenance work that an open-weight deployment forces onto the buyer. The premium is real work someone else does for you. The counter to the counter: a growing tier of inference providers — Groq, Together, Fireworks — bundle that productization around open-weight models and narrow the gap. How fast that tier matures is part of what decides the outcome.
So the moat is really a pricing-power question. If Anthropic and OpenAI can keep their prices high enough to fund the capex, but low enough that enterprises don’t bother wiring up their own open-weight stack, the ~$1.8 trillion mark holds. If open-weight tooling matures faster than that band allows, the mark compresses. Both outcomes are legitimate. The concentration risk is that the current mark leaves very little room for the second one.
5. Concentration #2 — Exclusive Training Data
The concentration in data is not that training data exists — training data is everywhere. It is that exclusive premium licensing deals are consolidating what remains legally acquirable into a handful of publishers negotiating with a handful of closed labs.
This is where the substrate war becomes ammunition. Closed labs have kept their lead partly through what they train on: books, articles, code, forum threads, images, and web pages that open-weight teams cannot easily replicate. If premium training data consolidates into a handful of exclusive deals with the closed labs, the capability gap widens and the $1.8 trillion valuation looks defensible. If open-weight teams find their way to comparable data — synthetic training runs, permissive corpora, favorable court rulings, new licensing structures — the gap narrows and the valuation looks vulnerable.
One clean distinction is worth naming, because the word “data” gets used two very different ways in AI conversations:
Training data is what a lab uses to build a model — the books, articles, images, code, and web pages ingested during pre-training. This is where the concentration sits, and it is the ammunition in the substrate war.
Customer data is what an enterprise sends through the model at inference time — prompts, documents, patient records, deal memos, transaction histories. This is not concentrated; it lives inside every enterprise that uses AI. The governance question around it — whether providers can prove, beyond contracts and attestations, that customer data does not flow into future training runs — is a live and important one. It is a lengthy post for another day.
On the training-data side, a shrinking pool of suppliers now controls most of what a frontier lab can legally acquire — Reddit, News Corp, the Financial Times, Stack Overflow, Shutterstock, and a handful of others. Every future model runs on inputs licensed from those suppliers, and the surviving frontier labs are signing exclusive deals to lock down what is left. That directly reinforces the closed-lab moat: exclusive premium training data is a substrate-war weapon, not a legal footnote.
How we got here matters. Through roughly 2022, the industry practice was to scrape the open web at enormous scale and argue fair use later. Common Crawl, Books3, LAION-5B — the datasets that trained the models that became today’s frontier — were assembled long before the legal rules were clear. The reckoning is now underway. A landmark 2025 court ruling drew the line explicitly: training on lawfully acquired books is fair use; training on pirated books is not. A $1.5 billion proposed settlement followed. Publisher suits against OpenAI and others are the next wave.
That reckoning helps the closed labs. New entrants cannot repeat the scrape-first playbook, licensing costs are rising, and the labs already at the frontier can afford exclusives that new entrants cannot. Dario Amodei has been out in front of the shift, pressing publicly for safety standards, model transparency, and licensing regimes. Read in good faith, an incumbent is helping the industry mature past a scrape-first era that had to be addressed and pushing back on distillation and piracy that would erode any lab’s moat. Read cynically — the case David Sacks and the All-In hosts have made explicitly — a safety-branded incumbent is using a national-security story to lobby for rules that raise the bar on cheaper open-weight rivals, exactly as those rivals catch up. Both readings can be true at once. What matters for the ~$1.8 trillion mark is which reading wins in Washington and state capitols over the next 18 months.
Regulation sets the calendar for the ammunition pipeline. State and federal AI laws now impose transparency and safety obligations on frontier models, and different levels of government have started challenging each other’s rules — which means that for buyers in regulated sectors, where a workflow runs is becoming as important as what it does. Every exclusive licensing deal, every court ruling, every new AI law makes the closed-lab moat wider or narrower. The ~$1.8 trillion mark in Concentration #1 depends on that pipeline holding.
Close
Public debates from the past few weeks have not just been about philosophy. It has been a contest over who owns the substrate of the next economy, and at which layer enterprise value settles. Today, there are ultimately two high stakes variables within frontier labs that influence destiny: model pricing power and exclusive training data.
That is the HyperDensity in the model and data substrate—subscribe here to receive a future issue in which we examine another part of the substrate wars: compute and power.
Cooper Westendarp is an investor, entrepreneur, and CEO Advisor at Reservoir Advisors. He writes HyperDensity on the world’s largest concentrations of capital, power, and risk. Find him on X at @cooperwest or @cooperwestendarp on LinkedIn.







