Great ActuatorThesis

Reference index — Sector 5: Data vendors, RL environments, evals, expert networks

Compiled 2026-07-14. Every figure below is sourced to a live link. Market caps carry an as-of date.


Mercor

What it is: The pace-setting neutral expert-data marketplace — matches PhDs, lawyers, bankers and physicians to AI labs for training and evaluation work; the thesis's proof that the "Switzerland" position captures rent while captive estates grow. Status: Private (San Francisco; founded 2023 by Brendan Foody, Adarsh Hiremath, Surya Midha). Investors include Felicis, Benchmark, General Catalyst. Market cap / valuation: $10B post-money at the ~$350M Series C (September/October 2025); in talks as of 2026-07-09 to raise ~$500M at a ~$20B valuation (TechCrunch and Forbes, 2026-07-09). The $20B mark is a reported negotiation, not a closed round. Latest financials: Crossed ~$2B annualized revenue run-rate in June 2026, roughly doubling from $1B in February 2026 and $760M at end-2025. Critical caveat the thesis relies on: this is gross customer spend before contractor payouts — contractors take roughly 60–70% of top line. Reported to be paying >$2M/day to its contractor network. Cited for: I.6 (neutral vendors compounding while Scale is captured; "$350M at $10B", "$2B annualized"); II.1 and II.2 (vendor revenue rotation, gross-vs-net caveat); II §"Vendor valuation air" (marks price a market ~10× audited demand); VI.4 (expert-knowledge collection is a real, fast-growing business — "~$2B annualized, >$2M/day to contractors"). ARGUMENT part-1 §"The external neutral market"; part-2 ARGUMENT §II table; part-6 ARGUMENT VI.4. Links:


Surge AI

What it is: The bootstrapped RLHF/expert-data house that quietly out-earned Scale — the second pace-setter of the neutral vendor position, and the thesis's cleanest example of capital-efficient rent capture in human data. Status: Private (San Francisco; founded 2020 by Edwin Chen, ex-Google/Twitter ML). Bootstrapped with zero outside VC until 2025; Chen still held ~75% of the company as of September 2025. Market cap / valuation: No closed priced round. Reported to be in talks (from July 2025) to raise ~$1B primary+secondary at a valuation of at least $25B (Bloomberg, 2025-07-30); earlier reporting the same month put the ask at >$15B. Treat as a talks-range mark of ~$15–25B, unclosed as of 2026-07-14. Latest financials: ~$1.2B revenue in 2024 — above Scale AI's $870M that year — reportedly profitable, on ~110–130 full-time employees plus a contractor/annotator network of roughly 1M registered (~50k active experts). Customers include OpenAI, Google, Microsoft, Meta and Anthropic. Cited for: I.6 ("Surge AI — bootstrapped, ~121 employees; revenue ~$1.2B, surpassing Scale's"; Meta's own TBD Labs researchers reportedly preferring Surge and Mercor); II.2 ("Surge ~$1.2B (Dec 2024) → ~$1.4B (Aug 2025)"); VI.4 ("bootstrapped, ~110 employees, seeking ~$15B on >$1B trailing revenue"). ARGUMENT part-1 §2 and §"Vendor valuation air" in part-2. Links:


Scale AI (and the Meta stake)

What it is: The data-labeling incumbent that sold neutrality — the thesis's natural experiment: when a vendor ties itself to one lab, its rivals take the business. Status: Private. Meta Platforms holds ~49% (non-voting) after a ~$14.3B investment announced 2025-06-12/13; co-founder/CEO Alexandr Wang left to run Meta Superintelligence Labs, with Jason Droege taking over as CEO. Market cap / valuation: ~$29B implied post-money at the Meta deal (June 2025); prior mark $13.8B (May 2024, Accel-led Series F). No newer priced mark found as of 2026-07-14. Latest financials: $870M revenue in 2024, exiting the year at a ~$1.5B annualized run rate, with pre-deal projections above $2B for 2025. Post-deal, Google (reportedly a ~$200M/yr customer) and OpenAI moved work away; Scale laid off 14% of staff (~200 people) in July 2025 and pivoted toward enterprise/government (its "Scale for Government" and applied AI lines). Cited for: I.6 — the whole sub-chapter: "$14.3B for 49% (June 2025), implied ~$29B", Google's ~$200M planned annual spend pulled, the subsequent layoffs, and the rotation of demand to Surge/Mercor/Turing/Handshake/Invisible/micro1. II.2 ("Scale AI post-Meta"). Also the Google–Appen precedent line (I.6 flagged items). Links:


Invisible Technologies

What it is: Ops-plus-AI-data house (custom training data, annotation, "AI process outsourcing") — one of the neutral vendors the thesis names as a beneficiary of Scale's capture. Status: Private (New York / fully remote; founded 2015 by Francis Pedraza). Market cap / valuation: >$2B, at a $100M round led by Vanara Capital announced 2025-09-16 (Acrew, Greycroft and others participating); total capital raised ~$144M. Latest financials: ~$134M revenue in 2024, up ~123% from ~$60M in 2023 (a >48× increase 2020→2024). Company claims it has trained foundation models for >80% of leading model providers (customers named include Cohere, Microsoft, AWS). Cited for: II.2 ("Invisible Technologies — ops/AI-data; ~$134M revenue 2024; $100M raise at >$2B"); I.6 (named among the neutral beneficiaries of the Scale–Meta rotation); II §"Every vendor revenue/valuation datum" verification list. Links:


Turing

What it is: Coding-data and engineering-talent supplier to OpenAI and other frontier labs — the "code RL data" leg of the neutral vendor set. Status: Private (Palo Alto; founded 2018 by Jonathan Siddharth and Vijay Krishnan). Market cap / valuation: $2.2B post-money at a $111M Series E announced 2025-03-06, led by Khazanah Nasional (Malaysia's sovereign fund) with AltaIR, WestBridge and Sozo; ~$225M raised to date. Latest financials: ~$300M ARR at the Series E (March 2025), up from $167M when the round was priced; the company describes itself as profitable. Cited for: II.2 ("Turing — coding-data/eng supplier to OpenAI and other labs; ~$300M ARR, $2.2B"); I.6 (neutral beneficiary list); II §"Every vendor revenue/valuation datum" verification list ("Turing $300M ARR / $2.2B"). Links:


micro1

What it is: The fastest-growing of the challenger data vendors and — the point the thesis leans on — the largest *visible commercial* egocentric/physical-capture program: it ships kits (including Meta Ray-Ban glasses) to contributors who record everyday manipulation tasks to build a robotics pre-training dataset. Status: Private (San Francisco; founded 2022/2023 by Ali Ansari, now 25). Market cap / valuation: $500M post-money at a $35M Series A led by 01 Advisors, announced 2025-09-12. Later reporting (Forbes, 2025-12-04) describes the company as being valued in the multibillions on secondary/new-round interest, but no priced round above $500M is confirmed to a primary source as of 2026-07-14. Latest financials: ARR path: ~$7M (start of 2025) → $50M (Sept 2025) → crossed $100M ARR (announced 2025-12-04)>$200M ARR (early 2026). Sacra estimates ~$300M annualized in April 2026, up from ~$125M at end-2025. Contributor pay for the egocentric program is reported around ~$15/hr. Cited for: II.2 ("micro1 — largest visible commercial egocentric/expert program; ~$300M ARR"); II.2 table ("micro1 ~$125M → ~$300M ARR (Apr 2026)"); II §"honest small end" — the thesis's caution that the "~90% of AI spend on real-world data" line is micro1's own CEO describing his own book (MIT Technology Review); II.3 wage-ladder line (micro1 contributors ~$15/hr); I.6 (neutral beneficiary list). Caution for the thesis: the $300M ARR figure traces to Sacra's estimate, not a company statement; the hard company-confirmed numbers are $100M ARR (Dec 2025) and >$200M ARR (early 2026). Links:


Snorkel AI

What it is: Programmatic-labeling pioneer turned expert-data and *evaluator* vendor — the thesis's example of the labeling layer migrating up into evals ("Snorkel Evaluate", "Expert Data-as-a-Service"). Status: Private (Redwood City; founded 2019 out of the Stanford AI Lab by Alex Ratner, Chris Ré and co-founders). Market cap / valuation: $1.3B post-money at a $100M Series D led by Addition, announced 2025-05-29 (Prosperity7, Greylock, Lightspeed, BNY, QBE Ventures participating). Total raised ~$237M. This is a ~30% step-up from its $1B mark in 2021. Latest financials: Snorkel does not disclose revenue; no audited or company-confirmed revenue figure exists as of 2026-07-14. The Series D announcement discloses only the raise, the $1.3B valuation, the $237M cumulative funding and the launch of Snorkel Evaluate and Expert Data-as-a-Service. Cited for: II.2 ("Snorkel AI — programmatic labeling → expert data + evaluators; $100M Series D at $1.3B"); II §verification list, which flags "a Snorkel discrepancy" — the flagged discrepancy resolves as: $1.3B is the Series D valuation (May 2025), not a revenue figure, and the older $1B figure is the 2021 valuation. Links:


Handshake (Handshake AI)

What it is: The college-recruiting network that turned its expert supply into a data-labeling business — and, in the thesis, the buyer that emptied the "neutral data-quality auditor" slot by acqui-hiring Cleanlab. Status: Private (San Francisco; founded 2013 by Garrett Lord, Ben Christensen, Scott Ringwelski). Market cap / valuation: ~$3.3B (reported at the time of the Cleanlab acquisition, January 2026). Its last publicly priced round was a $200M Series F at $3.5B (2022); the $3.3B figure is the mark reported in current coverage. Latest financials: Handshake AI (the data arm, launched ~2025) is reported to serve eight major AI labs including OpenAI, with projected "high hundreds of millions" of ARR for 2026. No audited figure is public. Cited for: II.2 ("Handshake — labeling vendor that acqui-hired Cleanlab"); II §"the neutral-auditor slot is structurally vacant and got emptier (Cleanlab acqui-hired by Handshake Jan 2026)"; I.6 (neutral beneficiary list). Note for the thesis: the Cleanlab deal was an acqui-hire of ~9 people including co-founders Curtis Northcutt, Jonas Mueller and Anish Athalye; terms were not disclosed. Cleanlab had raised ~$30M (Menlo Ventures, Bain Capital Ventures). Links:


Mechanize

What it is: The purest RL-environment startup — builds a small number of high-fidelity simulated "digital office" environments and evals for frontier coding agents and sells them to labs; in the thesis, the emblem of the new auto-gradable-task data market. Status: Private (San Francisco; founded 2025-04-17 by ex-Epoch AI researchers Tamay Besiroglu, Matthew Barnett and Ege Erdil). Market cap / valuation: No disclosed valuation. Total funding ~$9.1M (angel/seed; investors include Nat Friedman, Daniel Gross and Patrick Collison). Use funding, not valuation — none has been published as of 2026-07-14. Latest financials: No revenue disclosed. The company is pre-disclosure; its public statements describe selling environments/evals to leading AI labs. Cited for: II.1 ("Mechanize (founded Apr 2025, ~$9.1M raised, builds a small number of high-fidelity RL environments)"); II ARGUMENT §"RL environments (Mechanize founded Apr 2025…)" as evidence that the environments market is new, real and tiny relative to the data market. Links:


Prime Intellect

What it is: The open RL-environment stack — the Environments Hub (community-contributed RL environments), the Verifiers library, prime-rl training framework, plus hosted RL post-training, evals and on-demand GPU; the thesis's counter-example of environments as an open commons rather than a proprietary chokepoint. Status: Private (San Francisco / Berlin; founded 2024 by Vincent Weisser and Johannes Hagemann). Market cap / valuation: $1B post-money at a $130M round announced 2026-07-08, led by NVIDIA's NVentures with Intel Capital and Dell Technologies Capital (reported variously as Series A/Series B). Prior: seed/Series A rounds in 2024–25 (Founders Fund-led $15M). Latest financials: Annualized revenue topped ~$100M ahead of the July 2026 round; ~6,000 customers on the platform. Environments Hub hosts 2,500+ community-contributed environments — the number the thesis cites. Cited for: II.1 ("Prime Intellect — open Environments Hub, 2,500+ community environments"); II ARGUMENT §"RL environments … Prime Intellect's 2,500+ community environments" as evidence the environments layer is commoditizing in the open rather than concentrating. Links:


LMArena

What it is: The human-preference leaderboard (formerly LMSYS Chatbot Arena, out of UC Berkeley) commercialized into an evaluation business — the thesis's single hard commercial datapoint that a neutral, non-lab eval institution is fundable. Status: Private (Berkeley/San Francisco; spun out of LMSYS by Anastasios Angelopoulos, Wei-Lin Chiang and Ion Stoica). Market cap / valuation: $1.7B post-money at a $150M Series A announced 2026-01-06, led by Felicis and UC Investments with a16z, Kleiner Perkins and Lightspeed. That is ~3× the $600M valuation of its $100M seed (May 2025). Latest financials: ~$30M annualized "consumption rate" as of December 2025, under four months after launching its commercial product. No audited revenue. Cited for: II.4 ("LMArena — the eval business, commercialized: human-preference leaderboard → $1.7B / $150M Series A / >$30M run-rate"); II ARGUMENT §"a neutral, non-lab eval institution is fundable — LMArena — $150M Series A at $1.7B"; and II §"LMArena run-rate is the only hard commercial datapoint found" in the eval-market sizing. Links:


METR (Model Evaluation & Threat Research)

What it is: The independent, non-commercial evaluations institute — the thesis's example of an independent evaluation institution that exists but does *not* sell attestation as a business. Status: Private non-profit — a US 501(c)(3) research institute in Berkeley, CA (spun out of ARC Evals, led by Beth Barnes). Market cap / valuation: Not applicable — non-profit, no equity. Funding: philanthropic grants, historically led by Open Philanthropy; The Audacious Project catalysed ~$38M for "Canary" (METR + RAND), of which ~$17M supports work at METR (announced 2024-10-09). Latest financials: No revenue in the commercial sense; METR does pre-deployment evaluations for frontier labs (OpenAI o3/o4-mini/GPT-4.5, Anthropic Claude models) and publishes system-card contributions. Funding is grant-based; the ~$17M Audacious allocation is the largest disclosed commitment. Cited for: II.4 ("METR — the independent evaluation institution, non-commercial model"); II ARGUMENT §"None is a commercial neutral owner selling the attestation as a business." Also the widely used task-completion time-horizon result: the length of software tasks AI agents can complete at 50% reliability has doubled roughly every 7 months over ~6 years — METR's headline metric, relevant wherever the thesis reasons about capability trend lines. Links:


Epoch AI

What it is: The research nonprofit whose data-stock and scaling forecasts underwrite the thesis's entire "are we running out of text?" section — the source of the ~300T-token public-human-text figure and of the walked-back "data wall" window. Status: Private nonprofit research institute (founded 2022; director Jaime Sevilla). Funded by grants and donations, plus contractual research relationships; it publishes its funders. Market cap / valuation: Not applicable — nonprofit, no equity or valuation. Funding is disclosed on its transparency page rather than as a raise. Latest financials: No commercial revenue reported; Epoch discloses grant/donation funding and contract work on its transparency page (funders have included Open Philanthropy and, for specific benchmarks like FrontierMath, contracted work for OpenAI). Cited for: II.1 and II ARGUMENT §"Base: Epoch's effective stock of quality-and-repetition-adjusted public human text ≈ 300T tokens (90% CI 100T–1,000T)"; the "models trained on datasets roughly equal to the stock of public human text between 2026 and 2032" window (Villalobos et al., arXiv:2211.04325, last revised 2024-06-04); the thesis's claim that Epoch moved the front of the window out (~2026 → ~2028) and puts a scaling slowdown by 2040 at ~20%; the ~2.5×/yr algorithmic-efficiency figure (Ho et al. 2024); and repetition to ~4 epochs being near-free. Also cited in II.1's vendor-market model. Links:


Harvey

What it is: The legal-AI company built on applied professional rules and firm workflows — the thesis's proof that a corpus of *applied rules* (as distinct from published law) has monetizable value. Status: Private (San Francisco; founded 2022 by Winston Weinberg and Gabriel Pereyra). Backed by OpenAI Startup Fund, Sequoia, Kleiner Perkins, GIC, a16z. Market cap / valuation: $11B post-money at a $200M growth round co-led by GIC and Sequoia, announced 2026-03-25 — up from $8B in December 2025. Total raised >$1B. Latest financials: ~$300M ARR estimated as of May 2026 (Sacra), up from ~$195M at end-2025. Company-disclosed adoption: 142,000+ lawyers across 1,500+ customers in 60+ countries, including ~50% of the Am Law 100. Cited for: II.5 ("Harvey ~$300M ARR / $11B" as evidence the applied-rules corpus monetizes; "the enforcement-in-practice slot is empty" partially resolved by legal AI); II ARGUMENT §"Unlock: the applied-rules corpus has monetizable value. Harvey (~$300M ARR, $11B, May 2026)". Note the ARR figure is a Sacra estimate, not company-disclosed; the $11B/$200M round is primary. Links:


Thomson Reuters (TRI)

What it is: The incumbent owner of the professional-rules corpus (Westlaw, Practical Law) that folded generative AI (CoCounsel) into its core product — in the thesis, the counterweight showing that the incumbent corpus-holder, not the model, captures the applied-rules rent. Status: Public (NASDAQ: TRI; also TSX: TRI). Controlled by The Woodbridge Company (the Thomson family). Market cap / valuation: ~$41.1B as of 2026-07-14 (share price ~$93.32), per StockAnalysis. Latest financials: Q1 2026 (reported May 2026): revenue $1.924B, +8% organic; adjusted EBITDA $881M (42.2% margin); adjusted EPS $1.23 (vs $1.12 in Q1 2025); free cash flow $332M (+19% y/y). Legal Professionals segment revenue $756M, +9% organic, driven by Westlaw Advantage and CoCounsel. FY2026 guidance reaffirmed at 7.5–8% organic growth, margin ~40%. Cited for: II.5 ("Thomson Reuters folded CoCounsel into Westlaw so every seat gets it" — the incumbent-capture claim); II ARGUMENT §"legal AI (Harvey ~$300M ARR / $11B; Thomson Reuters CoCounsel…)" as the verified July-2026 evidence that the applied-rules layer monetizes. Links:


GLG (Gerson Lehrman Group)

What it is: The world's largest expert network — the pre-AI incumbent of paid expert access, and the thesis's benchmark for how big the expert-knowledge market was *before* the labs started buying it. Status: Private (New York; founded 1998 by Mark Gerson and Thomas Lehrman). PE-backed (Silver Lake, Bessemer, SFW Capital). Filed for a US IPO in October 2021 and withdrew it in March 2022 — so no public mark and no audited public financials since. Market cap / valuation: No current valuation is public. The last hard reference point is the withdrawn 2021/22 IPO filing; GLG has never disclosed a post-withdrawal mark. Could not source a current valuation to a primary link as of 2026-07-14. Latest financials: ~$650M revenue in 2021 (the figure from the IPO-era filing period, the most recent primary-adjacent number). Aggregator trackers put current revenue at ~$634M. Scale: >1M registered experts / freelance consultants, ~3,000 full-time staff, offices in 22 cities. Cited for: II.3 ("GLG: world's largest, >1M experts, revenue ~$634M"); II ARGUMENT §"a durable ~$3B market (2025), projected >$4.86B by end-2026 (GLG ~$634M…)". The thesis's own flag — "conflicting figures ($400M vs $634M); pin a year" — resolves as: $650M is the 2021 filing-period figure; $634M is an aggregator estimate with no stated year and should be labelled as such; the ~$400M figure could not be sourced to any primary document. Market context (same citation): Inex One sizes the expert-network industry at ~$3B in 2025, growing ~12%/yr, with projections above $4.86B by end-2026; ~11,200 firms use expert networks in 2026. Links:


Appen (ASX: APX)

What it is: The listed, pre-LLM-era crowd data-annotation incumbent — and the thesis's precedent case: when Google (its largest customer) walked, the equity collapsed. It is the "Google–Appen" half of the corroborating precedent for the Scale–Meta neutrality law. Status: Public (ASX: APX), Sydney-headquartered. Market cap / valuation: A$239.0M as of 2026-07-14 (share price A$0.89), per StockAnalysis — down ~20% y/y and roughly two orders of magnitude below its 2020 peak (~A$4.3B). Latest financials: FY2025: revenue US$233.4M, flat on FY2024; net loss US$21.8M (loss widened ~9%). Ex-Google revenue grew 4.5% with margin improvement and generative-AI-driven growth (notably in China). FY2026 guidance: revenue US$270–300M, 5–10% EBITDA margin. Google terminated its contract with Appen in January 2024 (~A$83M of FY2023 revenue). Cited for: I.6 flagged-precedent line — "(Google–Appen; OpenAI/Google–Scale)" as the corroborating precedent that a captured/dependent vendor loses its business; the thesis suggests demoting the Mercor-breach episode and using Google–Appen instead. Also the general "generic labeling commoditizes" claim in II.2. Links:


Innodata (NASDAQ: INOD)

What it is: The listed pure-play AI data-engineering vendor — the only publicly audited window onto what Big Tech actually pays for training data, and therefore the thesis's reality check on private vendors' unaudited ARR claims. Status: Public (NASDAQ: INOD), Ridgefield Park, NJ. Market cap / valuation: ~$2.2B as of 2026-07-14 (share price ~$67.51), per StockAnalysis. Latest financials: Q1 2026 (reported 2026-05-07): revenue $90.1M, +54% y/y; adjusted gross margin 47%; adjusted EBITDA $25.0M (28% of revenue); net income $14.9M ($0.46 basic EPS); cash $117.4M. Raised FY2026 revenue growth guidance to ~40%+, and disclosed new Big Tech engagements expected to contribute ~$51M in 2026. Cited for: II.2's audited-demand argument — the thesis's claim that private "vendor valuation air" prices a market ~10× above any audited demand rests on comparing Mercor/Surge marks with the only *audited* AI-data revenue in the sector, of which Innodata is the cleanest listed example (~$90M/quarter, growing ~50%). Links:


TELUS Digital

What it is: The large-scale CX-and-data-services arm whose "AI Data Solutions" line (annotation, training-data development, human-in-the-loop) is the enterprise, non-frontier end of the labeling market — a comparator showing where generic annotation revenue ended up: inside a telecom, not a lab supplier. Status: No longer public. Wholly owned subsidiary of TELUS Corporation (NYSE/TSX: T) since 2025-10-31, when TELUS completed the take-private of the ~43% it did not already own. Previously listed as NYSE/TSX: TIXT. Market cap / valuation: Take-private priced at US$4.50/share, aggregate consideration ~US$539M for the minority, an enterprise value of roughly US$2.9B including debt (announced June 2025, closed 2025-10-31). No standalone market cap exists after that date; the value now sits inside TELUS Corp. Latest financials: TELUS Digital no longer reports separately; TELUS guided to ~$150M/yr of operational efficiencies from the integration. Its last full year as a public company (FY2024) produced roughly US$2.65B of revenue — after which the AI Data Solutions line was folded into TELUS's reporting. Cited for: II.2/II.3 background on the commoditization of generic annotation — the segment of the market that does *not* command frontier-lab pricing and consolidates into incumbents. TELUS Digital is not named by figure in the thesis text; it is the comparator that supports the "generic capture commoditizes" claim in VI.4. Links:


Sama

What it is: The impact-sourcing annotation company (Kenya, Uganda, India delivery) — the low-wage end of the data-labour ladder the thesis prices against expert data, and the counter-example of labeling as a social-enterprise business rather than a frontier-lab supplier. Status: Private (San Francisco; founded 2008 by Leila Janah as Samasource; CEO Wendy Gonzalez). Originally a nonprofit, now a B-corp-style for-profit with a nonprofit foundation. Market cap / valuation: No disclosed valuation. Total funding ~$84.8M, of which a $70M Series B (2021) — reported at the time as the largest round for a female-led AI infrastructure company. The IFC has separately considered an investment of up to US$10M to expand Sama's Uganda operations. Use funding, not valuation: Sama has never disclosed a post-money valuation. Latest financials: Sama does not publish revenue; no audited figure is available (aggregator estimates circulate but are unsourced). Company-disclosed impact metrics: >65,000 people moved out of poverty through its training/employment model, validated by an MIT-led randomized controlled trial; customers include ~25% of the Fortune 50 (GM, Ford, Microsoft, Google named). Cited for: II.3's wage-ladder argument (the gap between commodity annotation labour — India gig work at ~$1/hr — and expert data at ~$85+/hr on Mercor); Sama is the named, documented instance of the impact-sourcing labour tier. Could not source Sama revenue to a primary link as of 2026-07-14 — only funding and impact metrics are disclosed. Links:


Labelbox

What it is: The labeling-tooling company that turned into a data factory plus expert marketplace (Alignerr) — the tooling-to-services migration that the thesis treats as the standard path in this market. Status: Private (San Francisco; founded 2018 by Manu Sharma and Brian Rieger). Operates the Alignerr expert network as its labour supply and Alignerr Connect as a hiring marketplace. Market cap / valuation: ~$1B at a $110M Series E led by SoftBank Vision Fund 2 and Andreessen Horowitz (October 2025); ~$189M raised across five rounds. Latest financials: No company-disclosed revenue. Third-party trackers put revenue around ~$115M/yr — an estimate, not audited, and it should be flagged as such if used. Cited for: Not named by figure in the thesis text; it is a comparable for II.2's vendor set — the tooling layer (Labelbox, Encord) commanding ~$0.5–1B marks versus the expert-supply layer (Mercor, Surge) commanding $10–25B, which is the size gradient the thesis's "vendor valuation air" argument depends on. Links:


Encord

What it is: The self-described "data layer for physical AI" — multimodal data curation/annotation for robotics, humanoids and AVs; the closest listed-in-thesis analogue to a tooling business aimed squarely at the robot-data problem. Status: Private (London / San Francisco; founded 2020 by Eric Landau and Ulrik Stig Hansen; Y Combinator). Market cap / valuation: ~$550M post-money at a $60M / €50M Series C led by Wellington Management (announced 2026), following a $30M Series B led by Next47 (August 2024). Total raised ~$110M (~€93M). Latest financials: No revenue disclosed. Company-disclosed operating metrics at the Series C: >5 petabytes of multimodal customer data under management and ~10× revenue growth from physical-AI customers over the prior year. Cited for: Not named by figure in the thesis text; supports II.2 and Part VI's claim that the physical/robot-data tooling market is real but an order of magnitude smaller than the expert-LLM-data market — a $550M tooling mark against Mercor's $20B talks. Links:


Lightwheel

What it is: Simulation-plus-data infrastructure for embodied AI — SimReady assets, synthetic and egocentric human data, and robot evaluation platforms; the clearest counterpart on the physical side to what Mercor is on the knowledge side, and evidence that the robot-data market has begun to price. Status: Private (Shanghai/San Francisco; founded 2023). Investors include Sequoia, Andreessen Horowitz, Lux Capital, Spark Capital and (in a later round) Ant Group. Market cap / valuation: Valued "north of $600M" after a $145M Series B led by Sequoia (with a16z, Lux, Spark); a further round led by Ant Group was reported in May 2026, directed at data and evaluation infrastructure. No exact post-money for the Ant round is public. Latest financials: ~$100M in Q1 2026 orders across simulation, data generation, evaluation and deployment (company announcement). Disclosed data scale: >1.5M hours of human data delivered, covering 25,000+ environment nodes and 100,000+ task types. Cited for: II.2 and Part VI's physical-capture argument — the thesis's claim that western physical-capture businesses barely exist while the collection market is real (best-evidence real-data spend ~$150–400M/yr mid-2026). Lightwheel's ~$100M of Q1-2026 orders is the strongest single datapoint that embodied-data collection is now a priced market, and it is not a western company — which sharpens the thesis's "expert-data yes, western physical-capture no" line in VI.4. Links:


Sunday Robotics

What it is: Home-robot company whose data strategy is the point — its Skill Capture Glove collects human manipulation demonstrations directly, bypassing teleoperation; the thesis's live instance of "collect the dark matter yourself rather than buy it." Status: Private (San Francisco / Palo Alto; founded 2024 by Tony Zhao and Cheng Chi, the ALOHA/diffusion-policy researchers). Market cap / valuation: $1.15B post-money at a $165M Series B led by Coatue, announced 2026-03-12 (Tiger Global, Benchmark, Bain Capital Ventures participating). Latest financials: Pre-revenue on the product: its Memo household robot begins beta deliveries in late 2026 (targeted by Thanksgiving 2026). Cost disclosure: prototypes ~$20,000 per unit, with a stated target below $10,000 at scale. No revenue reported. Cited for: Part VI's collection argument and II.2's physical/egocentric data line — the vertically integrated alternative to buying egocentric data from micro1 or Lightwheel. It is the "capture-glove, not teleop" instance behind the thesis's claim that the scarce end of the data market is being solved in-house by robot companies rather than sold by vendors. Links:


DoorDash (NASDAQ: DASH) — the Tasks program

What it is: The largest existing gig workforce turned into a real-world data-capture network: Tasks, launched 2026-03-19, pays Dashers to film everyday activities (loading a dishwasher, folding clothes), record unscripted speech in other languages, and shoot location photography — the thesis's proof that a consumer logistics platform can enter the physical-data collection business with distribution nobody else has. Status: Public (NASDAQ: DASH). Market cap / valuation: ~$82.4B as of 2026-07-14 (share price ~$189.17), per StockAnalysis. Latest financials: Q1 2026 (reported 2026-05-06): revenue $4.04B (+~33% y/y); marketplace GOV $31.6B (+37%); total orders 933M; adjusted EBITDA $754M; net income $184M (−5% y/y); operating margin 3.7%. Tasks is not separately monetized or broken out; DoorDash discloses that Dashers have completed more than 2 million tasks since 2024, so the program predates the standalone app. Cited for: Part VI's collection argument (VI.4 "Collection businesses… as generic capture commoditizes — now") and II.2's real-world/egocentric data line: DoorDash Tasks is the incumbent-distribution counterweight to micro1's kit-shipping program, and it constrains the thesis's "western physical-capture doesn't exist" claim — it now does, inside a $82B logistics company, using ~8M couriers. Footage is used to evaluate DoorDash's own models and partners' models in retail, insurance, hospitality and tech. Not available in California, New York City, Seattle or Colorado. Links:


Cross-cutting notes for the thesis

  • The Mercor gross-vs-net caveat is real and load-bearing. Mercor's "$2B annualized" is *gross customer spend*; contractors take ~60–70%. Net revenue to Mercor is therefore roughly $600M–$800M annualized. Any comparison of Mercor's $2B to Innodata's audited ~$360M annualized (Q1 2026 × 4) must state which basis it uses.
  • Only two audited windows exist in this sector: Innodata (NASDAQ: INOD, ~$90.1M revenue in Q1 2026) and Appen (ASX: APX, US$233.4M FY2025, loss-making). Everything else — Mercor, Surge, micro1, Handshake, Harvey — is self-reported ARR or a third-party estimate (Sacra). The thesis's "vendor valuation air" argument stands on exactly this asymmetry.
  • The valuation gradient by layer (all marks as of mid-2026): expert supply (Surge $15–25B talks, Mercor $20B talks, Handshake ~$3.3B, Turing $2.2B, Invisible >$2B) ≫ tooling (Labelbox ~$1B, Encord ~$550M) ≈ environments (Prime Intellect $1B, Mechanize ~$9.1M raised) ≈ evals (LMArena $1.7B). Physical capture prices lower and mostly outside the West (Lightwheel >$600M).

Highlights