Part II · 13 min
Part II: The Missing Data
Most of what people know was never online.
1. Running out of internet
300Teffective stock of public textGPT-4 can code just about anything because GitHub and Stack Overflow were sitting on the open internet, free to take. Same story for essays, contracts, translations: whatever the web had piles of, the models got good at. That trick has a measured ceiling. Epoch AI's estimate puts the effective stock of quality-adjusted public human text around 300 trillion tokens, with a 90% confidence band from 100 trillion to a quadrillion, and an exhaustion window between 2026 and 2032. Epoch's own later work widens the frame (counting every quality tier and modality, the effective band stretches to the quadrillions), and its own window has already slipped from 2026 to 2028. No single exhaustion year is credible. Algorithmic efficiency improves about 2.5× per year, repetition to four epochs is nearly free, and multimodal transcription keeps extending the base. The wall moves.
What does not move is the direction of the spending. Whatever the exact ceiling, spend is rotating out of the drained pools and up toward data that has to be commissioned, and the rotation is visible in every price on the board.
$340 → $118an hour of remote-piloted demonstrationThe cheap end is flooding. An hour of remote-piloted robot demonstration fell from about $340 in early 2024 to about $118 by March 2026. Simple egocentric video runs $15–30 an hour. Off-the-shelf SLAM and hand-pose pipelines now compute the kinematic labels vendors used to sell, and free fully-annotated corpora (Ropedia's Xperience-10M shipped ten thousand hours in March 2026) put a labeled floor under the whole market.
The pool that bought the field time is now itself a paid market. Reinforcement learning on auto-gradable tasks, math and code, extended the free ride, and then minted its own vendor class: Mechanize builds a small number of high-fidelity coding-agent environments for frontier labs; Prime Intellect runs an open hub with 2,500-plus community environments. The Information reported Anthropic leadership discussing over $1 billion on RL environments in a single year; exclusive environment deals reportedly price at four to five times non-exclusive. When the free extension of a resource has to be bought, the resource is functionally scarce whatever the token counter says.
~$2BMercor annualized revenue, grossAnd the commissioned end is compounding faster than any published market model. Grand View Research sizes data collection and labeling at $3.77 billion in 2024 growing to $17.1 billion by 2030, a 28% annual clip from a named analyst. The vendors are outrunning it. Mercor crossed roughly $2 billion in annualized revenue in June 2026, doubling in four months, with the caveat that the figure is gross customer spend before contractor payouts, and contractors take 60–70% of it. The efficiency gains the skeptic cites cut the other way here: every 2.5× improvement in tokens-per-FLOP raises the value of each marginal high-quality token, which is why the commissioned expert hour gets more expensive while the scraped hour gets cheaper.
~$26TARK full-substitution ceilingThe downstream market this data gates is forecast anywhere from $29.5 billion (IDTechEx's realizable case) to ARK's ~$26 trillion full-substitution ceiling, conditional on humanoids operating at scale. The data market is small next to the capability it feeds. That mismatch creates both the opportunity and the risk.
2. What was never written down
The bigger pool was never online. Most of what people know isn't written down anywhere: how a therapist runs a session, how the one technician fixes the line at 3am, how a caseworker actually applies a rule, how a carpenter picks his cut. It sits in people's heads and inside companies. The models can't learn it because nobody recorded it. That is the dark matter of the economy, and by value it is most of the economy.
Collecting it is becoming a market with two ledgers that disagree.
$1.4BSurge AI revenue, ~50k contractorsThe vendor ledger is spectacular. Surge AI bootstrapped to roughly $1.2 billion of revenue by December 2024 and ~$1.4 billion by August 2025 on about 120 employees and fifty thousand contractors, and opened its first raise at a reported $15–25 billion, still unclosed. Mercor runs at ~$2 billion gross. Turing reports ~$300 million ARR at a $2.2 billion valuation, profitable. Invisible did ~$134 million in 2024 and raised at over $2 billion. micro1 confirmed $100 million ARR in December 2025 and over $200 million in early 2026 (the ~$300 million figure circulating is a Sacra estimate), and it runs the largest visible egocentric program: four thousand contributors in 71 countries recording 160,000-plus hours a month at ~$15 an hour. Snorkel repositioned around expert data and evaluators at a $1.3 billion valuation; it discloses no revenue at all. Scale, post-Meta, guides toward ~$2 billion while its former customers' spend feeds everyone above.
$150–400Mexternal paid egocentric and manipulation dataThe demand ledger is thinner than the valuations imply. The widely quoted "$100 million-plus a year spent on real-world data" is one vendor's CEO describing his own book. No independently audited figure exists for the physical-data market anywhere. The estimate used here of external paid egocentric and manipulation data lands at $150–400 million a year as of mid-2026, up from under $50 million in 2024: real, fast, and an order of magnitude below what the valuations price. The buyer list is the dozen names from the previous part, four of whose most aggressive programs collect in-house. Vendor revenue splits by segment (expert LLM data versus coding and RL versus physical capture versus evaluation) are undisclosed at nearly every one of these companies, so any exact figure for the physical slice is an estimate, including this one.
~$3B/yrexpert-network industry, total revenueExpert networks provide a priced precedent for standalone human know-how. The industry (GLG at ~$650 million in its 2021 IPO-filing period, AlphaSights above $300 million, Third Bridge around $287 million) runs about $3 billion a year in total, heading toward $4.9 billion by end-2026. That is the mature market price of un-instrumented human know-how, before any copy-forever multiplier from turning the conversation into training data. GLG's own history carries the warning label: its share of that industry fell from roughly half to a quarter over a decade as the category fragmented. Collection businesses commoditize, and even the winner's seat fragmented.
Capture costs remain high. A skilled teleoperator produces 30–80 usable demonstrations a day on a $50,000–150,000 rig; a manipulation policy wants 300–1,200 clean demonstrations to generalize; the expert whose judgment you actually want bills $65–120 an hour and cannot be swapped for a $22-an-hour generalist without losing the thing being captured. The market is clearing anyway: enterprise buyers now treat $50,000–150,000 per task as an in-reach data budget, a threshold that was out of reach for most organizations two years ago. The asset arithmetic explains why: the expert's hour used to be worth her wage, and a recorded expert hour is copied into every model and every downstream worker forever. The wage stops being the ceiling.
Where the objection genuinely holds: tacit judgment. For the highest-value dark matter (the caseworker's discretion, the integration of a career), no one has shown a capture method whose cost per useful hour clears, and it may only ever be collected as the byproduct of machines working alongside people. That question stays open.
Rare-case coverage, verified labels, clean rights, and independent tests kept out of training retain value as raw collection floods.
3. Rules
Statute text is online. How a rule actually gets applied is not: the discretion, the exceptions, the local practice, the difference between what the regulation says and what the inspector accepts. That gap is exactly what an agent operating in the world needs, and nobody has organized its collection at scale.
$11BHarvey valuation, March 2026Legal AI provides the nearest commercial proof that applied-rules corpora are valuable. Harvey raised $200 million at $11 billion in March 2026 (its ~$300 million ARR is an analyst estimate; the round is confirmed) selling models tuned on firms' own documents to 100,000-plus lawyers. Thomson Reuters folded CoCounsel into Westlaw so that every answer cites into its proprietary applied-law corpus; the moat is precisely the record of how law gets used. Both monetize applied law. Neither collects enforcement-as-applied for the physical and operational world: how the food-safety inspector actually scores a kitchen, how the building examiner actually reads the code. The value is proven one domain over. The collection stands unorganized.
$8.2Bnational AI industry fundMeanwhile the rules about the teaching itself are being written, and one country is writing them first. On February 28, 2026, China's MIIT released a 52-standard system for humanoid robots and embodied intelligence through its technical committee TC08, built with more than 120 institutes and companies across six domains, from components to safety and ethics. Inside it sit the referee slots for data specifically: a technical requirement for embodied-intelligence data-generation platforms at call-for-comment stage, a data-quality specification in drafting, a high-quality-dataset standard in pre-research, a collection specification under approval. A secondary April 2026 count puts China's state-backed collection centers at at least ninety across 23 provinces, but no primary source has been identified; the confirmed 2025 baseline was forty-plus. An $8.2 billion national fund and provincial programs finance the buildout. The rules-writer and the biggest customer are the same actor.
The Western side of the ledger is nearly empty. ISO/IEC's 5259 series covers data quality for machine learning generically, with no embodiment scope. SGS issued the world's first certification against it in November 2025; the installed base is approximately one. NVIDIA's Halos program took the machine-safety certification seat in June 2026 with an accredited inspection lab and six certification bodies, and no dataset scope. The one ISO humanoid-dataset work item is pen-held by Chinese institutions, with no certification mechanism defined. Europe's dataset-quality harmonized standard is pre-draft.
Two dates convert this from committee work into gatekeeping. On January 20, 2027, the EU Machinery Regulation applies: self-certification for machine-learning safety components ends and notified bodies inherit conformity assessment they currently lack tooling for. In August 2028, the AI Act's high-risk provisions give notified bodies documentation access to robot-embedded AI. The definitions of "qualified" (qualified dataset, qualified training process) decide what may travel to market. Whoever writes the rules of teaching ends up refereeing the whole market, and there is a precedent for how completely that seat can capture an industry: a shared file of who paid their bills, started by two brothers in Atlanta in 1899, became Equifax, and today a conforming US mortgage cannot be originated without pulling the bureaus' records, because the secondary market that finances the loan requires it. The collected record didn't just inform the industry. It became the gate the industry has to clear, installed by statute and underwriting rules, which is exactly what the 2027 dates are the beginning of here.
4. People
It runs backward too. Once the best practitioner's know-how is captured, everyone else trains on it, human and machine, forever. A worker's recorded know-how can be worth more than the worker's output. Today the entire capture economy pays for time (India gig collectors around a dollar an hour, China's centers around three, micro1's contributors around fifteen), and the lesson rides along unpriced.
The machinery for pricing it already exists, because entertainment built it first and recently, under strike pressure.
~10 wordsone AI-voice line, paid per unitSAG-AFTRA's 2023 film and television agreement established the consent architecture: creating or using a digital replica of a performer requires separate, conspicuous, reasonably specific written consent, and the consent is void if the use exceeds the description the performer signed. A captured likeness is licensed for a stated purpose. The 2025 Interactive Media agreement, ratified in July after a year-long strike, added the compensation mechanics: AI-generated voice work is paid per line (a line runs about ten words), and performers can suspend consent for new AI-generated material during a strike. Machine output derived from a person's captured self now has a per-unit price and a revocation right, in a signed national contract. The writers got the ownership terms: under the 2023 WGA agreement, AI output cannot count as source material against a writer's credit, a company cannot require a writer to use AI, and AI-generated material handed to a writer must be disclosed.
The courts are building the floor underneath. In Lehrman v. Lovo, voice actors alleged a text-to-speech company obtained recordings under an "internal research" pretense and cloned them; in July 2025 the Southern District of New York let the right-of-publicity claims proceed, including under New York's new digital-replica provision, and dismissed the copyright claims, since a voice as such can't be copyrighted. The emerging shape: the actionable wrong is unconsented commercial use of a person's captured identity, which is precisely the terrain of occupational know-how capture. The consent has to be real, too. Courts and regulators in Kenya and the Philippines have already invalidated token-induced consent for biometric capture; consent bought with a gratuity is not consent. Clean, documented, jurisdiction-valid rights are an asset. A competitor's dataset with dirty rights is a legal liability, whatever it is valued at.
Outside entertainment, no union contract or employment agreement yet covers know-how capture or teleoperation. Warehouse and logistics labor has fought automation for jobs; nobody has yet negotiated the price of the lesson. The template is written and tested one industry over. The person teaching the system that replaces him should be getting paid for the lesson, and the first contracts that say so will look a lot like a per-line voice clause with a hard hat on.
5. The exams
When everything needs to be collected, the scarce skill is knowing how to build the collection: tasks, environments, and above all the tests that grade the results. An independent test kept out of training is the strongest asset in the stack because the buyer cannot make one alone. A lab that writes its own evaluation can train against it, turning the test into a measure of memory rather than generalization. Grading has to come from outside, making it the one layer the captive estates of Part I cannot in-source.
29%of MMLU items showing contamination signsContamination is measured. A Johns Hopkins analysis found roughly 29% of MMLU test items showing contamination signs; swapping contaminated items for clean equivalents dropped model scores by up to 13 points; a 31-model study found the standard math benchmarks leaking into training data broadly, with leakage growing over time. A public benchmark is a depreciating asset that decays on contact with the thing it measures. A secret, continuously refreshed one is durable, which is an economic statement, and the market has started pricing it.
$150MLMArena raise at $1.7B valuationMETR runs frontier evaluations on a donation-funded model and takes no money from the AI companies it grades: neutrality as constitution. ARC-AGI keeps handcrafted environments secret by construction; on ARC-AGI-3, humans score near 100% and frontier models near 0.5%, with over $2 million in prizes. LMArena proved the commercial case: from a preference leaderboard to paid evaluation services, $150 million at a $1.7 billion valuation in January 2026, with an annualized run-rate above $30 million within four months of launching its first product. Neutral grading is a fundable business now.
For the physical world (the hardest version, since an embodied exam has to run on hardware), the owners today are academic: RoboArena's distributed double-blind benchmark across DROID platforms, RoboDojo's 42 simulated and 18 real tasks with a live leaderboard. No commercial neutral owner of an independent physical-world testing program exists. The adjacent history says the seat doesn't fill itself: the autonomous-vehicle data era, 2017 to 2022, ran on identical friction and produced no standalone data-grading company; the flagship independent label-QA startup, Cleanlab, was acqui-hired in January 2026. The seat has been getting emptier while the need compounds.
$1.09BETS revenue, fifty million testsMature testing businesses show what the position can be worth. ETS, a nonprofit, administers more than fifty million tests a year in 180-plus countries on about $1.09 billion of revenue, and everyone upstream (students, schools, employers, immigration systems) organizes around its attestation. UL Solutions runs $3.05 billion of revenue at a 25.9% EBITDA margin. The testing-inspection-certification industry does roughly $254 billion a year at 18–23% EBITDA. At the top of the stack sit the ratings agencies: S&P's ratings segment at a ~63% operating margin, Moody's around 60%, exam-owners whose exam is a letter grade on a bond. The pattern goes back as far as collected records do: one mortality table, compiled for the Society for Equitable Assurances in the 1760s, turned life insurance from gambling into a priceable industry and stood as the standard for a century. A single credible dataset founded the industry that priced against it.
The AI version is harder in one respect: the graded party is an adversary that will train on any leaked item, so the exam needs continuous fresh production in a way a mortality table never did. It is also why the position, once held, is so defensible. Almost nobody owns one yet. For embodied models, nobody does. You can't scrape the physical world; you have to go get it, and going to get it takes parts, deployments, and money.