Where to buy AI training data in 2026

By , founder of grokkedPublished Updated 7 min read

Frontier labs and neolabs buy training data from five places: brokers that license enterprise workflow data from companies, expert human-data vendors such as Scale AI, Surge AI and Mercor, licensing deals with publishers and platforms, data marketplaces, and the public web plus synthetic data. For real workplace data, the chat, email, tickets, documents and code that show how companies actually operate, the source is a broker that licenses it from the companies themselves.

grokked is that broker. It licenses de-identified workplace data from operating companies across 12 sectors, with signed supplier agreements, GDPR anonymization and AI Act provenance documentation. Labs preview samples in the catalog before they license.

Key takeaways

  • Enterprise workflow data is the scarcest input: Google bid $10 million for Spirit Airlines’ emails, Teams chats and files in 2026, against $7.5 million from Mercor and a later $12.5 million offer from Micro1.
  • Expert-data vendors are now billion-dollar businesses. Mercor passed $2 billion in gross annualized revenue in 2026; Surge AI made more than $1 billion a year by 2025.
  • Publisher deals buy public-facing content, at $16 million to $60 million a year for the largest. They don’t show how companies work inside.
  • The public web is shared and running out, and indiscriminate training on model-generated data causes “model collapse”.
  • Under the EU AI Act, general-purpose model providers must summarize their training data, including data from intermediaries. Buy with provenance documentation.

Why sourcing matters more for neolabs

Radical Ventures defines a neolab as “a startup focused on long-term technical breakthroughs, typically founded by research scientists and engineers from leading AI industrial or academic organizations”, and counts 40+ neolabs that raised $40 billion over three years (Radical Ventures). The Decoder describes them as specialized AI startups founded by researchers who left major AI companies (The Decoder).

NeolabReported fundingSource
Safe Superintelligence$2B at a $32B valuation (2025)TechCrunch
Thinking Machines Lab$2B seed at a $12B valuation (2025); in talks at $40B+ (2026)TechCrunch, PYMNTS
Reflection AI$2B at an $8B valuation (2025)TechCrunch
Periodic Labs$300M seed (2025)Maginative

Money buys compute, but everyone can buy compute. NEA expects the winners to be “research-first companies that own a domain-biased corpus” (NEA). The public web won’t provide that: Common Crawl’s 300 billion+ pages are free to anyone (Common Crawl), and Epoch AI expects the stock of public human text to be fully used between 2026 and 2032 (Epoch AI). The data you can license and others can’t is part of the edge.

The five places labs buy data

SourceWhat you getHow buying worksExamples
Enterprise workflow data brokersDe-identified chat, email, tickets, documents, code and records from operating companiesPreview samples in a catalog, then licensegrokked; also Troveo, SimpleClosure, Ooak Data
Expert human data and RL environmentsPreference data, expert tasks, evaluations, RL environmentsCustom projects, sales-ledScale AI, Surge AI, Mercor, Handshake, Micro1
Publisher and platform licensesNews, forum, image and book archivesBilateral licensing dealsReddit, News Corp, Shutterstock, Dotdash Meredith
Data marketplacesThird-party datasets of every kindBrowse, sample, then license or subscribeDefined.ai, Datarade, AWS Data Exchange
Public web and synthetic dataCrawled pages; model-generated dataFree download, or generated in-houseCommon Crawl

Enterprise workflow data

Agents need to see work being done. Fortune put it plainly: “an agent needs to learn how to choose actions, use tools, respond to intermediate results and recover from mistakes” (Fortune). The labs say the same about their sourcing. OpenAI wants datasets “not already easily accessible online to the public today” (OpenAI), and Anthropic lists “non-public datasets obtained from businesses” among its training sources (Anthropic).

Prices show the demand. Google won the bankruptcy auction for Spirit Airlines’ internal data at $10 million, ahead of Mercor at $7.5 million (Forbes), and Micro1 later offered $12.5 million (Fortune). RL-environment builders want the same material: TechCrunch reported Anthropic leaders discussing more than $1 billion of spending on RL environments in a year (TechCrunch), and Deeptune raised a $43 million Series A for environments that simulate the work of “accountants, customer support reps, and DevOps engineers” (Fortune).

Bankruptcy auctions are one-offs. grokked gives labs a steady supply from operating companies:

  • Workflow-rich data. Messages, documents and tickets that show how professional work gets done, across support, sales and CRM, code and engineering, procedures, contracts, chat and email, finance and projects.
  • 12 sectors. Accounting, engineering, logistics, software, insurance and more, from small firms to large enterprises.
  • Clean license chain. Every supplier signs an agreement and a GDPR data processing agreement. Personal and business-confidential details are removed before licensing, and every dataset ships with the provenance documentation the AI Act requires.
  • Clear license terms. Licensed for training and evaluating models, for at most 3 years, followed by certified deletion. No re-identification, resale or verbatim reproduction.
  • Preview first. Labs see a sector and a size band, never the supplier’s name, and preview de-identified samples before they license.

License real workplace data.

Preview de-identified samples from operating companies, with provenance documentation.

Request access

Expert human data and RL environments

Expert-data vendors produce data to order: you specify the tasks, they recruit the people. The category has grown into one of the biggest in AI:

  • Scale AI. Meta invested $14.3 billion for a 49% stake in 2025 (CNBC).
  • Surge AI. Annual revenue above $1 billion without outside funding, as of mid-2025 (SiliconANGLE). It now also sells frontier data and RL environments “off the shelf” (Surge AI).
  • Mercor, Handshake and Micro1. Mercor hit $2 billion in gross annualized revenue in summer 2026 and Handshake $1 billion earlier that year, while Micro1 grew from $100 million to $500 million in eight months (TechCrunch).

Expert vendors are the right choice for targeted tasks, rubrics and evaluations. What they can’t easily produce is years of real operating history from inside a company, which is what enterprise workflow data adds.

Publisher and platform licenses

The largest content deals license public-facing archives:

SellerBuyerReported priceSource
RedditGoogleAbout $60M a yearReuters
News CorpMetaUp to $50M a year, for three yearsEngadget
The New York TimesAmazon$20M–$25M a yearTheWrap
Dotdash MeredithOpenAIAt least $16M a yearEngadget
ShutterstockAI companies$104M in AI licensing revenue in 2023Bloomberg Law

Per item, Defined.ai’s CEO told Reuters in 2024 that buyers pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films, and that “the market rate for text is $0.001 per word” (Reuters). These deals cover what an organization publishes, not how it works.

Data marketplaces

Marketplaces let you browse and sample before you license. Defined.ai calls itself “the world’s largest AI data marketplace” (Defined.ai); Datarade lists datasets from registered businesses, including AI training data (Datarade); on AWS Data Exchange, providers may not list data that identifies a person unless it is publicly available (AWS). Human Native AI, a UK marketplace for creator content, joined Cloudflare in January 2026 (Cloudflare).

Quality and provenance vary by seller, so ask each one for its supplier agreement and anonymization method.

Public web and synthetic data

Common Crawl maintains “a free, open repository of web crawl data that can be used by anyone” (Common Crawl). Free isn’t rights-free: the Commission’s AI Act template notes that “the public availability of the datasets for free does not mean that the content at issue is necessarily free of rights” (European Commission).

Synthetic data has limits too. A 2024 Nature paper found that “indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear” (Nature). Real human data keeps those tails.

Due diligence before you license

  1. Provenance. Who supplied the data, under what agreement, and can the vendor show it?
  2. Personal data. Is it anonymized to the GDPR standard, where no one is identifiable by means reasonably likely to be used (GDPR, recital 26)? Pseudonymized isn’t enough.
  3. AI Act summary. Since August 2, 2025, general-purpose model providers must publish a training-content summary (AI Act). The template separates data licensed from rightsholders from datasets obtained through “data intermediaries” (European Commission). Ask which bucket the data falls in and for documents you can reuse.
  4. How it was obtained. Acquisition matters as much as use: Anthropic agreed to pay $1.5 billion, roughly $3,000 per book, and to destroy datasets in the authors’ class action over pirated books (CNBC).
  5. License scope. Training and evaluation rights, term, deletion at the end, and limits on reproduction.
  6. Samples first. Preview real records before you commit.

License real workplace data.

Preview de-identified samples from operating companies, with provenance documentation.

Request access

Questions

Where can AI labs buy training data?

From five places: brokers that license enterprise workflow data from companies, such as grokked; expert human-data vendors such as Scale AI, Surge AI and Mercor; publisher and platform licenses; data marketplaces such as Defined.ai, Datarade and AWS Data Exchange; and the public web, mainly Common Crawl.

What is a neolab?

A research-first AI startup, usually founded by scientists and engineers from leading labs. Radical Ventures counts 40+ neolabs that raised $40 billion over three years, including Safe Superintelligence, Thinking Machines Lab and Reflection AI (Radical Ventures).

How much does AI training data cost?

It ranges from fractions of a cent per word for generic text to $1–$2 per image (Reuters), and from tens of millions of dollars a year for large publisher archives to $10 million for one airline’s internal emails, chats and files (Forbes).

Is licensed enterprise data compatible with the EU AI Act?

Yes, if it comes with provenance. General-purpose model providers must summarize their training content, including private datasets from data intermediaries (European Commission). grokked supplies the provenance documentation with every dataset.

Can we preview grokked data before licensing?

Yes. Labs preview de-identified samples in the catalog and see each dataset’s sector and size band before licensing. Request the catalog.

See what your data is worth.

Get an estimate in 60 seconds. Listing your data costs nothing.

Get paid in 7 days

Sources

  1. Radical Ventures, June 21, 2026. The rise of NeoLabs
  2. The Decoder, March 14, 2026. Ex-Anthropic researchers launch AI startup Mirendil to tackle scientific research
  3. TechCrunch, April 12, 2025. OpenAI co-founder Ilya Sutskever’s Safe Superintelligence reportedly valued at $32B
  4. TechCrunch, July 15, 2025. Mira Murati’s Thinking Machines Lab is worth $12B in seed round
  5. PYMNTS, September 3, 2026. Thinking Machines Lab seeks $1 billion at $40 billion valuation
  6. TechCrunch, October 9, 2025. Reflection AI raises $2B to be America’s open frontier AI lab, challenging DeepSeek
  7. Maginative, September 30, 2025. Periodic Labs launches with $300M to build an “AI scientist”
  8. NEA, June 17, 2026. The AI neolab wild west
  9. Common Crawl. Common Crawl
  10. Epoch AI, June 6, 2024. Will we run out of data to train large language models?
  11. Fortune, September 14, 2026. Little-known AI startup Micro1 tries to trump Google’s bid for bankrupt Spirit Airlines’ data
  12. OpenAI, November 9, 2023. OpenAI Data Partnerships
  13. Anthropic. Non-user privacy policy
  14. Forbes, August 19, 2026. AI companies desperate for data are buying up dead airlines’ emails and scanning old books
  15. TechCrunch, September 16, 2025. Silicon Valley bets big on ‘environments’ to train AI agents
  16. Fortune, March 19, 2026. Andreessen Horowitz backs Deeptune’s $43M Series A to build ‘training gyms’ for AI agents
  17. CNBC, June 12, 2025. Scale AI founder Wang announces exit for Meta part of $14 billion deal
  18. SiliconANGLE, July 1, 2025. Data labeling startup Surge AI reportedly seeking $1B in first capital raise
  19. Surge AI. Surge AI
  20. TechCrunch, August 20, 2026. AI data startup Micro1 reaches $500M gross run rate amid AI training boom
  21. Reuters, February 22, 2024. Reddit in AI content licensing deal with Google, sources say
  22. Engadget, March 3, 2026. Meta signs a multimillion dollar AI licensing deal with News Corp
  23. TheWrap, July 30, 2025. New York Times seals $20 million AI deal with Amazon
  24. Engadget, November 19, 2024. OpenAI will pay Dotdash Meredith at least $16 million per year to license its content
  25. Bloomberg Law, June 4, 2024. Shutterstock’s AI-licensing business generated $104 million
  26. Reuters, April 5, 2024. Inside Big Tech’s underground race to buy AI training data
  27. Defined.ai. Partnership Programs
  28. Datarade. Become a data provider
  29. Amazon Web Services. Publishing guidelines for AWS Data Exchange
  30. Cloudflare, January 15, 2026. Cloudflare strengthens content offering to AI companies with acquisition of Human Native
  31. European Commission, July 24, 2025. Explanatory notice and template for the public summary of training content for general-purpose AI models
  32. Nature, July 24, 2024. AI models collapse when trained on recursively generated data
  33. EUR-Lex, April 27, 2016. Regulation (EU) 2016/679 (General Data Protection Regulation)
  34. EUR-Lex, June 13, 2024. Regulation (EU) 2024/1689 (Artificial Intelligence Act)
  35. CNBC, September 5, 2025. Anthropic to pay $1.5 billion to settle authors’ copyright lawsuit

About the author

Nico Vergauwen

Nico Vergauwen is the founder of grokked, a data broker that licenses companies’ de-identified internal data to AI labs. Owners never pay anything and are paid within 7 days of each sale.

More guides