Where to buy AI training data in 2026
By Nico Vergauwen, founder of grokkedPublished Updated 7 min read
Frontier labs and neolabs buy training data from five places: brokers that license enterprise workflow data from companies, expert human-data vendors such as Scale AI, Surge AI and Mercor, licensing deals with publishers and platforms, data marketplaces, and the public web plus synthetic data. For real workplace data, the chat, email, tickets, documents and code that show how companies actually operate, the source is a broker that licenses it from the companies themselves.
grokked is that broker. It licenses de-identified workplace data from operating companies across 12 sectors, with signed supplier agreements, GDPR anonymization and AI Act provenance documentation. Labs preview samples in the catalog before they license.
Key takeaways
- Enterprise workflow data is the scarcest input: Google bid $10 million for Spirit Airlines’ emails, Teams chats and files in 2026, against $7.5 million from Mercor and a later $12.5 million offer from Micro1.
- Expert-data vendors are now billion-dollar businesses. Mercor passed $2 billion in gross annualized revenue in 2026; Surge AI made more than $1 billion a year by 2025.
- Publisher deals buy public-facing content, at $16 million to $60 million a year for the largest. They don’t show how companies work inside.
- The public web is shared and running out, and indiscriminate training on model-generated data causes “model collapse”.
- Under the EU AI Act, general-purpose model providers must summarize their training data, including data from intermediaries. Buy with provenance documentation.
Why sourcing matters more for neolabs
Radical Ventures defines a neolab as “a startup focused on long-term technical breakthroughs, typically founded by research scientists and engineers from leading AI industrial or academic organizations”, and counts 40+ neolabs that raised $40 billion over three years (Radical Ventures). The Decoder describes them as specialized AI startups founded by researchers who left major AI companies (The Decoder).
| Neolab | Reported funding | Source |
|---|---|---|
| Safe Superintelligence | $2B at a $32B valuation (2025) | TechCrunch |
| Thinking Machines Lab | $2B seed at a $12B valuation (2025); in talks at $40B+ (2026) | TechCrunch, PYMNTS |
| Reflection AI | $2B at an $8B valuation (2025) | TechCrunch |
| Periodic Labs | $300M seed (2025) | Maginative |
Money buys compute, but everyone can buy compute. NEA expects the winners to be “research-first companies that own a domain-biased corpus” (NEA). The public web won’t provide that: Common Crawl’s 300 billion+ pages are free to anyone (Common Crawl), and Epoch AI expects the stock of public human text to be fully used between 2026 and 2032 (Epoch AI). The data you can license and others can’t is part of the edge.
The five places labs buy data
| Source | What you get | How buying works | Examples |
|---|---|---|---|
| Enterprise workflow data brokers | De-identified chat, email, tickets, documents, code and records from operating companies | Preview samples in a catalog, then license | grokked; also Troveo, SimpleClosure, Ooak Data |
| Expert human data and RL environments | Preference data, expert tasks, evaluations, RL environments | Custom projects, sales-led | Scale AI, Surge AI, Mercor, Handshake, Micro1 |
| Publisher and platform licenses | News, forum, image and book archives | Bilateral licensing deals | Reddit, News Corp, Shutterstock, Dotdash Meredith |
| Data marketplaces | Third-party datasets of every kind | Browse, sample, then license or subscribe | Defined.ai, Datarade, AWS Data Exchange |
| Public web and synthetic data | Crawled pages; model-generated data | Free download, or generated in-house | Common Crawl |
Enterprise workflow data
Agents need to see work being done. Fortune put it plainly: “an agent needs to learn how to choose actions, use tools, respond to intermediate results and recover from mistakes” (Fortune). The labs say the same about their sourcing. OpenAI wants datasets “not already easily accessible online to the public today” (OpenAI), and Anthropic lists “non-public datasets obtained from businesses” among its training sources (Anthropic).
Prices show the demand. Google won the bankruptcy auction for Spirit Airlines’ internal data at $10 million, ahead of Mercor at $7.5 million (Forbes), and Micro1 later offered $12.5 million (Fortune). RL-environment builders want the same material: TechCrunch reported Anthropic leaders discussing more than $1 billion of spending on RL environments in a year (TechCrunch), and Deeptune raised a $43 million Series A for environments that simulate the work of “accountants, customer support reps, and DevOps engineers” (Fortune).
Bankruptcy auctions are one-offs. grokked gives labs a steady supply from operating companies:
- Workflow-rich data. Messages, documents and tickets that show how professional work gets done, across support, sales and CRM, code and engineering, procedures, contracts, chat and email, finance and projects.
- 12 sectors. Accounting, engineering, logistics, software, insurance and more, from small firms to large enterprises.
- Clean license chain. Every supplier signs an agreement and a GDPR data processing agreement. Personal and business-confidential details are removed before licensing, and every dataset ships with the provenance documentation the AI Act requires.
- Clear license terms. Licensed for training and evaluating models, for at most 3 years, followed by certified deletion. No re-identification, resale or verbatim reproduction.
- Preview first. Labs see a sector and a size band, never the supplier’s name, and preview de-identified samples before they license.
License real workplace data.
Preview de-identified samples from operating companies, with provenance documentation.
Expert human data and RL environments
Expert-data vendors produce data to order: you specify the tasks, they recruit the people. The category has grown into one of the biggest in AI:
- Scale AI. Meta invested $14.3 billion for a 49% stake in 2025 (CNBC).
- Surge AI. Annual revenue above $1 billion without outside funding, as of mid-2025 (SiliconANGLE). It now also sells frontier data and RL environments “off the shelf” (Surge AI).
- Mercor, Handshake and Micro1. Mercor hit $2 billion in gross annualized revenue in summer 2026 and Handshake $1 billion earlier that year, while Micro1 grew from $100 million to $500 million in eight months (TechCrunch).
Expert vendors are the right choice for targeted tasks, rubrics and evaluations. What they can’t easily produce is years of real operating history from inside a company, which is what enterprise workflow data adds.
Publisher and platform licenses
The largest content deals license public-facing archives:
| Seller | Buyer | Reported price | Source |
|---|---|---|---|
| About $60M a year | Reuters | ||
| News Corp | Meta | Up to $50M a year, for three years | Engadget |
| The New York Times | Amazon | $20M–$25M a year | TheWrap |
| Dotdash Meredith | OpenAI | At least $16M a year | Engadget |
| Shutterstock | AI companies | $104M in AI licensing revenue in 2023 | Bloomberg Law |
Per item, Defined.ai’s CEO told Reuters in 2024 that buyers pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films, and that “the market rate for text is $0.001 per word” (Reuters). These deals cover what an organization publishes, not how it works.
Data marketplaces
Marketplaces let you browse and sample before you license. Defined.ai calls itself “the world’s largest AI data marketplace” (Defined.ai); Datarade lists datasets from registered businesses, including AI training data (Datarade); on AWS Data Exchange, providers may not list data that identifies a person unless it is publicly available (AWS). Human Native AI, a UK marketplace for creator content, joined Cloudflare in January 2026 (Cloudflare).
Quality and provenance vary by seller, so ask each one for its supplier agreement and anonymization method.
Public web and synthetic data
Common Crawl maintains “a free, open repository of web crawl data that can be used by anyone” (Common Crawl). Free isn’t rights-free: the Commission’s AI Act template notes that “the public availability of the datasets for free does not mean that the content at issue is necessarily free of rights” (European Commission).
Synthetic data has limits too. A 2024 Nature paper found that “indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear” (Nature). Real human data keeps those tails.
Due diligence before you license
- Provenance. Who supplied the data, under what agreement, and can the vendor show it?
- Personal data. Is it anonymized to the GDPR standard, where no one is identifiable by means reasonably likely to be used (GDPR, recital 26)? Pseudonymized isn’t enough.
- AI Act summary. Since August 2, 2025, general-purpose model providers must publish a training-content summary (AI Act). The template separates data licensed from rightsholders from datasets obtained through “data intermediaries” (European Commission). Ask which bucket the data falls in and for documents you can reuse.
- How it was obtained. Acquisition matters as much as use: Anthropic agreed to pay $1.5 billion, roughly $3,000 per book, and to destroy datasets in the authors’ class action over pirated books (CNBC).
- License scope. Training and evaluation rights, term, deletion at the end, and limits on reproduction.
- Samples first. Preview real records before you commit.
License real workplace data.
Preview de-identified samples from operating companies, with provenance documentation.
Questions
Where can AI labs buy training data?
From five places: brokers that license enterprise workflow data from companies, such as grokked; expert human-data vendors such as Scale AI, Surge AI and Mercor; publisher and platform licenses; data marketplaces such as Defined.ai, Datarade and AWS Data Exchange; and the public web, mainly Common Crawl.
What is a neolab?
A research-first AI startup, usually founded by scientists and engineers from leading labs. Radical Ventures counts 40+ neolabs that raised $40 billion over three years, including Safe Superintelligence, Thinking Machines Lab and Reflection AI (Radical Ventures).
How much does AI training data cost?
It ranges from fractions of a cent per word for generic text to $1–$2 per image (Reuters), and from tens of millions of dollars a year for large publisher archives to $10 million for one airline’s internal emails, chats and files (Forbes).
Is licensed enterprise data compatible with the EU AI Act?
Yes, if it comes with provenance. General-purpose model providers must summarize their training content, including private datasets from data intermediaries (European Commission). grokked supplies the provenance documentation with every dataset.
Can we preview grokked data before licensing?
Yes. Labs preview de-identified samples in the catalog and see each dataset’s sector and size band before licensing. Request the catalog.
See what your data is worth.
Get an estimate in 60 seconds. Listing your data costs nothing.
Sources
- Radical Ventures, June 21, 2026. The rise of NeoLabs
- The Decoder, March 14, 2026. Ex-Anthropic researchers launch AI startup Mirendil to tackle scientific research
- TechCrunch, April 12, 2025. OpenAI co-founder Ilya Sutskever’s Safe Superintelligence reportedly valued at $32B
- TechCrunch, July 15, 2025. Mira Murati’s Thinking Machines Lab is worth $12B in seed round
- PYMNTS, September 3, 2026. Thinking Machines Lab seeks $1 billion at $40 billion valuation
- TechCrunch, October 9, 2025. Reflection AI raises $2B to be America’s open frontier AI lab, challenging DeepSeek
- Maginative, September 30, 2025. Periodic Labs launches with $300M to build an “AI scientist”
- NEA, June 17, 2026. The AI neolab wild west
- Common Crawl. Common Crawl
- Epoch AI, June 6, 2024. Will we run out of data to train large language models?
- Fortune, September 14, 2026. Little-known AI startup Micro1 tries to trump Google’s bid for bankrupt Spirit Airlines’ data
- OpenAI, November 9, 2023. OpenAI Data Partnerships
- Anthropic. Non-user privacy policy
- Forbes, August 19, 2026. AI companies desperate for data are buying up dead airlines’ emails and scanning old books
- TechCrunch, September 16, 2025. Silicon Valley bets big on ‘environments’ to train AI agents
- Fortune, March 19, 2026. Andreessen Horowitz backs Deeptune’s $43M Series A to build ‘training gyms’ for AI agents
- CNBC, June 12, 2025. Scale AI founder Wang announces exit for Meta part of $14 billion deal
- SiliconANGLE, July 1, 2025. Data labeling startup Surge AI reportedly seeking $1B in first capital raise
- Surge AI. Surge AI
- TechCrunch, August 20, 2026. AI data startup Micro1 reaches $500M gross run rate amid AI training boom
- Reuters, February 22, 2024. Reddit in AI content licensing deal with Google, sources say
- Engadget, March 3, 2026. Meta signs a multimillion dollar AI licensing deal with News Corp
- TheWrap, July 30, 2025. New York Times seals $20 million AI deal with Amazon
- Engadget, November 19, 2024. OpenAI will pay Dotdash Meredith at least $16 million per year to license its content
- Bloomberg Law, June 4, 2024. Shutterstock’s AI-licensing business generated $104 million
- Reuters, April 5, 2024. Inside Big Tech’s underground race to buy AI training data
- Defined.ai. Partnership Programs
- Datarade. Become a data provider
- Amazon Web Services. Publishing guidelines for AWS Data Exchange
- Cloudflare, January 15, 2026. Cloudflare strengthens content offering to AI companies with acquisition of Human Native
- European Commission, July 24, 2025. Explanatory notice and template for the public summary of training content for general-purpose AI models
- Nature, July 24, 2024. AI models collapse when trained on recursively generated data
- EUR-Lex, April 27, 2016. Regulation (EU) 2016/679 (General Data Protection Regulation)
- EUR-Lex, June 13, 2024. Regulation (EU) 2024/1689 (Artificial Intelligence Act)
- CNBC, September 5, 2025. Anthropic to pay $1.5 billion to settle authors’ copyright lawsuit
About the author
Nico Vergauwen
Nico Vergauwen is the founder of grokked, a data broker that licenses companies’ de-identified internal data to AI labs. Owners never pay anything and are paid within 7 days of each sale.