Part I · 14 min
Part I: The Center of Gravity
The labs pay whoever feeds the learning loop.
1. The money already moved
Microsoft will spend about $190 billion on datacenters in 2026. That is one company, one year, and it is roughly what the entire US federal government spends on highways, aviation, and rail combined. Amazon guided to about $200 billion, Alphabet to $180–190 billion, Meta to $125–145 billion. The four together guide to about $695–725 billion for the year, roughly 70–77% above about $410 billion in 2025, and approximately three-quarters of the guided total is AI infrastructure. On the earnings calls the operators describe the market as supply-constrained. Nobody guides to a 70%-plus increase on infrastructure they doubt.
$965BAnthropic closed, $47B run-rate revenueCapital. The private labs raise at prices that recently belonged to whole public industries. OpenAI closed a $122 billion round at an $852 billion valuation in March 2026. Anthropic closed $65 billion at $965 billion in May, on run-rate revenue that crossed $47 billion. xAI raised $20 billion at ~$230 billion, then merged into SpaceX and went public in June; the combined company opened around $2.1 trillion. Safe Superintelligence holds a ~$32 billion mark on ~$6 billion raised with no product. Mistral took on $830 million of debt in March 2026 to buy about 13,800 Nvidia chips: a lab borrowing against its future to buy compute, which tells you what the lab thinks the compute buys.
$8.5BA3 facility secured by GPU clustersCompute, and the debt behind it. Lenders invented a collateral class for GPUs and took it to investment grade in about thirty months. CoreWeave borrowed $2.3 billion against its chips in August 2023 and $7.5 billion in May 2024, when the structure was still exotic. By March 2026 its $8.5 billion facility carried an A3 from Moody's, secured by GPU clusters and contracted revenue including a reported ~$19 billion Meta backlog. Around it, 27 datacenter securitizations raised $13.3 billion in 2025, up 55% on the year, and total AI datacenter debt issuance passed $200 billion. Lenders are harder to convince than equity investors, and they are lending.
Talent. In mid-2025 Meta ran a hiring raid on frontier researchers with packages reported up to ~$300 million over four years. Ruoming Pang, who ran Apple's foundation models, reportedly moved for more than $200 million. One offer to a Thinking Machines co-founder was reported as high as $1.5 billion over six years, and was declined. Meta disputes some of the reported figures, and none of the offers is independently verified. Even with that qualification, frontier talent has concentrated around a few dozen training runs.
Capital, compute, talent: all three flows point at the same address, whoever trains the largest and best models.
The spending case has serious gaps. Sequoia's David Cahn has been asking where the revenue is since 2023; on his arithmetic the gap between what the buildout implies and what AI actually earns runs around $600 billion a year and has widened. The 2026 capex-to-AI-revenue ratio is roughly 10:1, where cloud in 2011 ran 2.4:1. Some of the demand is circular: Nvidia commits up to $100 billion to OpenAI, OpenAI signs ~$300 billion of cloud with Oracle, Oracle buys Nvidia chips to serve it, and Nvidia holds equity in most of the labs that buy from it. Michael Burry's version: the hyperscalers depreciate GPUs over five or six years when the economic life is closer to two or three, which flatters reported profit by an estimated $50–60 billion a year. And Daron Acemoglu's estimate of what AI adds to GDP, roughly 1.1–1.6% over a decade, is orders of magnitude below what the spending implicitly assumes.
$103.9BWorldCom assets, then-largest bankruptcyA crash could destroy the financing without reversing the centering. Carriers laid 80 million miles of fiber between 1996 and 2001 on roughly a trillion dollars of capital; four years after the bust, 85–95% of it was still dark; bandwidth prices fell about 90%; WorldCom filed the then-largest bankruptcy in American history in July 2002 with $103.9 billion of assets. The crash wiped out the financing while the glass stayed in the ground. Google started buying dark fiber around 2005, and on the cheap overbuilt network the next decade got YouTube, Netflix streaming, and AWS. The 1880s ran the same experiment with railroads at up to ~6% of GDP in capex: repeated panics, a quarter of US rail mileage in receivership after 1893, and the track consolidated under Morgan and carried the industrial economy for a century. In both cases the crash annihilated the capital structure and concentrated the asset. It did not reverse the centering.
The analogy breaks at asset life. Fiber lasts 25 years in the ground; a GPU is economically old at three. If the durable thing here is the silicon, a crash strands a rotting asset and the fiber precedent does not transfer. The durable layer would instead have to be the models, accumulated data, and deployment records, none of which depreciate on Nvidia's cycle. Whether those assets persist through a financing crash remains unresolved.
The open internet's training windfall is substantially spent. Thirty years of text, images, and code sat online, already written and paid for by everyone who posted. That windfall trained the current generation. The next input is the physical world, and the physical world does not come as a download.
2. Why the centering happens
An economy is mostly a machine for connecting information: which customer needs what, which worker can do which task, where capital should sit, what a thing is worth. Prices, firms, and middlemen have been the connectors for a couple of centuries. The models are becoming a better connector, and they improve monthly. Whoever owns the best one collects rent on coordination itself, which is a deeper tax base than any single product category. That is the prize, and it is why the training race consumes whatever it consumes.
This power compounds only while the model keeps learning. The loop runs model, then capital, then compute and data, then better model. A lab that stops learning starts decaying into a utility. For as long as the race runs, the labs are structurally enormous buyers of whatever feeds the loop, and they pay in the strongest currency in the economy. The condition holds today and there is no published date for when it stops.
NVIDIA's EgoScale work gives one measured view of the appetite. Task success scaled log-linearly with hours of action-labeled first-person video: 0.30 to 0.71 as the corpus went from one thousand hours to 20,854, an R² of 0.9983, with no saturation in sight. That curve supports continued buying.
Commoditization is already visible. DeepSeek's V4-Pro scores 80.6 on SWE-bench Verified against Claude's 80.8 at roughly 1/28 the price. API prices fell about 80% between early 2025 and early 2026; a million GPT-4-class tokens went from ~$60 to under $1. Open weights make frontier-adjacent capability nearly free to self-host. If the connector is a commodity, there is only a subsidy war, and OpenAI's reported ~$14 billion 2026 loss looks like the war's invoice.
$25BOpenAI revenue annualized by mid-2026The collapse in token prices did not collapse revenue. The market split into two tiers, commodity tokens near $0.14 per million and frontier reasoning tokens at $30–180, and the gap is widening. OpenAI's revenue reached roughly $25 billion annualized by mid-2026, about 70% of it subscriptions, with enterprise revenue growing from ~$1 billion annualized at the start of 2025 to more than $7 billion. The per-token price fell 80% underneath a subscription and enterprise business that septupled. Cheap inference turned out to be how the model inserts itself into more of the economy's coordination surface, the way 90%-cheaper bandwidth built the internet economy on top of the pipe. The rent gets collected where the model is embedded in a workflow, and that layer got more valuable as the input got cheaper.
Two conditions remain unresolved. If open weights let enterprises self-host and own their own coordination relationships, the rent exists but migrates to the enterprises; the labs' claim on it depends on stickiness that is suggested by the enterprise curve and proven by nothing yet. Revenue moving up-stack is a fact while profit is not: a $25 billion business under a $14 billion loss is collecting position ahead of rent. Whether the position converts is unknown. Every serious player is nevertheless paying as if it will.
3. The buyers
The serious buyers of physical-world training data number about a dozen worldwide, and the whole market fits in one contact list. Each runs a variant of the same purchase order (modest volumes of real, embodiment-matched demonstrations, large volumes of cheap egocentric human video, synthetic multiplication in between), and each is spending against the same unsaturated scaling curve. The current buyers and their approaches are:
NVIDIA. Sells the tooling and the models, consumes egocentric video at the largest disclosed scale. GR00T N1.7 pretrained on the 20,854-hour EgoScale corpus, which NVIDIA describes as beyond what teleoperation can scale to; Jim Fan's public roadmap talks about ten-million-hour datasets, which is roughly 480 times EgoScale. Its open PhysicalAI datasets run to 15 terabytes with over ten million downloads.
Tesla. The most captive program. In-house collection operators wearing five-camera helmet rigs at $25–48 an hour, factory-floor capture extended from Fremont to Austin, a reported thousand-plus Optimus units on live factory tasks by January 2026, and 2026 capex guided above $25 billion. The demand is captive and unavailable to an outside vendor.
Figure. Trained Helix v1 on about 500 teleop hours, then bought reach instead of hours: the Brookfield partnership opens 100,000 residential units and 500 million square feet of commercial space as a captive human-video estate, and navigation shipped trained on 100% human video with zero robot demonstrations. Series C above $1 billion at a reported $39 billion; the BotQ line was producing a robot every 90 minutes by April 2026, with roughly 740 operating by end of June.
1X. Turned the customer into the collection site. NEO sells at $20,000 or $499 a month, ships late 2026, and runs 60–70% autonomous at launch with remote operators covering the rest under documented privacy controls: no-go zones, face blur, opt-out. The buyer pays 1X to generate 1X's training data inside the buyer's home.
Physical Intelligence. About ten thousand hours of self-collected robot demonstrations, roughly $1.5 billion raised and deployed into collection infrastructure, and a documented buyer of third-party data on top. $600 million at $5.6 billion in November 2025; reported in talks at $11 billion in March 2026, which counts as talks until it closes.
Skild AI. Simulation-first on Isaac Lab and Cosmos, real data for post-training, one brain across many bodies with a cross-customer flywheel. Series C of about $1.4 billion at a reported $14–15 billion, roughly a tripling in seven months.
Google DeepMind. In-house ALOHA-2 teleoperation, plus Apptronik's Robot Park in Austin, a flagship collection facility expanded June 2026, feeding Apollo data into Gemini Robotics. Robot Park is DeepMind's captive capture node; the demand is one balance sheet's.
OpenAI. Restarted robotics with in-house collectors teleoperating Franka arms since February 2025, and a 2026 hiring list that reads like a data-factory org chart: actuator design, DAQ-station engineering, an operations manager for data acquisition. The stated near-term focus is robots that help build datacenter and grid infrastructure, which would make the robot program feed the compute program.
ByteDance. Seed Robotics trains GR-3 on egocentric VR-device teleoperation (reported as PICO hardware, though the technical report says only "VR devices"), plus the ByteMini 22-degree-of-freedom bimanual robot. Robotics spend is reported in the multi-billions; ByteDance files no audited financials, so no number can be pinned.
China's state-backed centers. A parallel system: more than forty government-backed collection centers, around two dozen operational, the Beijing site spanning ten thousand square meters across sixteen task categories, a Zigong facility rated at roughly three million data entries a year. AgiBot shipped 5,168 units in calendar 2025 by Omdia's count and passed its fifteen-thousandth cumulative robot in June 2026, with a million-plus trajectories published open. The system is essentially closed to Western vendors and sits on open-market pricing like a bulk-supply overhang.
Every serious addition to the list since 2025 is captive. Boston Dynamics launched electric Atlas production at CES 2026 with all 2026 units committed to Hyundai's Metaplant and Google DeepMind, inside Hyundai's ~$26 billion robotics commitment. Meta stood up a humanoid team under Marc Whitten with an Android-of-robotics licensing strategy and a glasses fleet (seven-million-plus units sold in 2025) as an ambient capture funnel. Samsung consolidated Rainbow Robotics for ~$181 million; Rainbow now trades around ₩13 trillion on KOSDAQ. Apple is the weak signal: a home hub in 2026 and a reported tabletop robot for 2027, an emerging buyer at most.
Demand has grown while its addressable share has shrunk, because the most aggressive buyers are building their own supply. The external market for raw physical-world data is small (the estimate used here runs $150–400 million a year, and the only outside anchor is one vendor CEO describing his own book), and it is being actively in-sourced by its richest customers. Selling bulk demonstrations into this list means selling a commodity to buyers who are building their own supply.
Captive estates cannot produce a neutral exam. A lab can collect its own training data, but a test set cannot stay independent if the same lab can inspect it and train against it, and a single-body captive estate cannot benchmark performance across rival platforms. The more the buyers self-supply, the more they need someone neutral holding the exam. Demand is concentrated among about a dozen names; losing a single buyer takes out close to a tenth of the market.
4. What happened to Scale
$14.3BMeta stake in Scale AIIn June 2025 the market ran the controlled experiment. Meta invested about $14.3 billion for a 49% non-voting stake in Scale AI, valuing it above $29 billion; founder Alexandr Wang left to lead Meta's superintelligence lab and kept a Scale board seat. Scale's work did not get worse that week. Its customer list evaporated anyway.
Google, Scale's largest customer with a reported ~$200 million planned for 2025, moved to cut ties within days. OpenAI confirmed a wind-down it had already begun. xAI and Microsoft backed away. By August Scale had laid off about 200 employees and 500 contractors and consolidated sixteen generative-AI teams into five. The buyers' logic was simple and uniform: a data vendor sees your roadmap, your prototypes, and your failure cases, and this one now answered to a competitor. Quality was never the question.
~$2BMercor revenue run-rate by JuneThe revenue moved to vendors that could still credibly claim to belong to no one. Surge AI, bootstrapped with about 121 employees, crossed roughly $1.2 billion in revenue, past Scale's reported $870 million, and opened its first raise, reported at $15–25 billion and still unclosed. Mercor, which markets itself explicitly as Switzerland, went from $350 million at a $10 billion valuation in October 2025 to a reported ~$2 billion revenue run-rate by June 2026, doubling in four months, with a raise reported in talks near $20 billion. Turing, Invisible, Handshake, and micro1 were all named among the gainers. Meta's own researchers reportedly preferred buying from Surge and Mercor over the Scale pipeline Meta now half-owned. That detail comes from a single-thread report and remains uncorroborated, though it is consistent with the broader episode. The pattern was not new either; Google had walked away from Appen in a comparable hurry in 2024. In this layer, buyer flight is the standard penalty, executed in days.
The rule the episode wrote down: sell to all of the learners and belong to none of them. The durable position is the layer every lab needs and no lab wants to build for its competitors.
Nvidia looks like a counterexample and is actually the boundary condition. Nvidia holds equity in nearly every lab it sells to (OpenAI, xAI, SSI, Figure, Skild, Physical Intelligence, Thinking Machines) and remains the universal supplier, unpunished. The difference is what the vendor can see. A chip performs identically for every buyer and reveals nothing about any of them. A data vendor reads the buyer's roadmap in every task spec it fulfills. Capture disqualifies vendors whose work gives them a window into the customer; it spares suppliers of inputs that carry no information back. That defines exactly who must stay neutral: anyone routing data, records, or evaluations between the labs.
Mercor adds a second constraint: security. In March 2026 it disclosed a supply-chain breach, a compromised CI/CD pipeline, exposing contractor data tied to lab research pipelines, and Meta paused all work with it indefinitely. The mechanism matches the Scale episode even though the failure does not: a shared layer carrying sensitive material is usable only while it is neutral and clean, and the penalty for compromise is immediate flight.