What company data AI labs pay for, and what it’s worth

By , founder of grokkedPublished Updated 6 min read

AI labs pay for data that shows how work actually gets done: team chat and email threads, support tickets, documents and contracts, code reviews, CRM pipelines, procedures and finance workflows. That data isn’t on the public web, which labs have largely used up. OpenAI says it wants datasets “not already easily accessible online to the public today” (OpenAI).

The prices are real but not public. Reported deals for one company’s workplace data run from $10,000 to several hundred thousand dollars (Gizmodo, Fortune), and Google won a bankruptcy auction for Spirit Airlines’ emails, Teams chats and files at $10 million in 2026. Below: the data types that sell, why private data is worth more, what labs have paid and how to estimate your own.

Key takeaways

  • The most valuable company data shows people solving real problems step by step: support tickets, code reviews and incidents, procedures, and the chat and email around them.
  • Public text is running out. Epoch AI puts the stock at about 300 trillion tokens and expects models to fully use it between 2026 and 2032.
  • Labs have paid about $60 million a year for Reddit’s data, a reported $250 million+ over five years for News Corp’s, and $5,000 per book in a HarperCollins deal.
  • What your data earns depends on data types, team size, years of history and industry. The grokked estimator gives a range in 60 seconds.

The data types AI labs buy

Labs train models to do professional work, so they want records of professional work. These are the eight types grokked licenses, with the tools they usually live in:

DataExamplesUsually inWhat a model learns from it
Customer supportTickets, chats, help-desk repliesZendesk, Freshdesk, IntercomHow real questions get diagnosed, answered and escalated
Sales & CRMPipelines, call notes, proposalsSalesforce, HubSpot, PipedriveHow deals are qualified, priced and negotiated
Code & engineeringRepositories, code reviews, incidentsGitHub, GitLab, JiraHow software is written, reviewed and fixed
Procedures & SOPsWikis, manuals, checklistsConfluence, Notion, SharePointStep-by-step methods for a trade
Contracts & documentsContracts, reports, bidsGoogle Drive, Word, DocuSignHow professional documents are structured and argued
Chat & emailTeam channels, email threadsSlack, Teams, OutlookHow decisions get made in conversation
Finance & accountingBookkeeping, invoices, ordersExact, SAP, OdooHow transactions are recorded, checked and closed
Projects & tasksPlans, tasks, status updatesAsana, monday.com, ClickUpHow work is planned and tracked over time

The Spirit Airlines auction shows the appetite. The dataset holds about 100 million emails, 500 million Teams chats, 17 million OneDrive files, more than 20 million SharePoint items and 516 code repositories, with personal information removed before transfer (TechSpot).

Workflow data matters most for AI agents. As Fortune put it, “an agent needs to learn how to choose actions, use tools, respond to intermediate results and recover from mistakes” (Fortune). A support ticket that goes from complaint to fix, or a code review that catches a bug, shows exactly that.

Why private data is worth more than public data

Most of the public web has already gone into training. Epoch AI estimates the stock of quality public human text at around 300 trillion tokens and expects models to fully use it between 2026 and 2032 (Epoch AI). “We’ve achieved peak data and there’ll be no more,” OpenAI co-founder Ilya Sutskever said at NeurIPS in December 2024. “There’s only one internet” (The Verge).

So labs look for data nobody else has. OpenAI’s data partnerships page says it is “particularly looking for data that expresses human intention (e.g., long-form writing or conversations rather than disconnected snippets)” (OpenAI). Anthropic’s privacy policy lists “non-public datasets obtained from businesses” among its training sources (Anthropic).

Forbes summed up the shift in August 2026: “Enterprise data has become incredibly useful for AI training, since labs had already hoovered up the entire public internet by late 2024” (Forbes).

What AI labs have paid for data

$60M / yr

Google’s reported license for Reddit’s data

Reuters, 2024

$250M+

OpenAI’s reported five-year deal with News Corp

TechCrunch, 2024

$5,000

per book in HarperCollins’ AI licensing deal, split with the author

Publishers Weekly, 2024

Those are the headline deals. Other disclosed prices:

SellerBuyerReported priceSource
RedditAI companies$203M in data licensing contracts over two to three yearsTechCrunch, from Reddit’s IPO filing
Taylor & Francis (Informa)Microsoft$10M+ access fee, plus recurring payments in 2025–2027Informa
ShutterstockAI companies$104M in AI licensing revenue in 2023Bloomberg Law
Dotdash MeredithOpenAIAt least $16M a yearEngadget
The New York TimesAmazon$20M–$25M a yearTheWrap
News CorpMetaUp to $50M a year, for three yearsEngadget
Spirit Airlines, in bankruptcyGoogle$10M at auction; Micro1 later offered $12.5MFortune
Startups closing downAI labs and data firms$10,000–$100,000 per companyGizmodo

Publisher deals pay for public-facing content. Company data is a newer market: per-company prices are reported, not listed, and “there is no reliable market average,” Fortune noted (Fortune).

What your company’s data could be worth

grokked’s estimator values each data type separately, then adjusts for team size, years of history and industry. Value grows with headcount, but less than proportionally, and we count the most recent 15 years. Data that shows step-by-step work, such as code, support tickets and procedures, carries the most training signal per person.

Example companyDataEstimated payout
20-person agency, 8 yearsChat & email, Contracts & documents, Projects & tasks€30K–€43K
40-person accounting firm, 20 yearsProcedures & SOPs, Contracts & documents, Chat & email€72K–€105K
60-person engineering firm, 15 yearsChat & email, Contracts & documents, Procedures & SOPs, Projects & tasks€155K–€225K
150-person software company, 10 yearsCode & engineering, Customer support, Chat & email, Procedures & SOPs€295K–€435K
500-person logistics company, 12 yearsCustomer support, Sales & CRM, Procedures & SOPs, Finance & accounting€605K–€885K

Indicative. The final amount depends on volume, quality and what buyers pay. For accountants, lawyers and healthcare, client files are excluded, and the estimate already accounts for that.

You’re paid per sale. AI labs preview de-identified samples in our catalog and license your dataset, and you’re paid within 7 days of each sale. You never pay anything; grokked takes a commission on each sale.

See what your data is worth.

Get an estimate in 60 seconds. Listing your data costs nothing.

Get paid in 7 days

What stays out

Some data never sells, whatever a buyer would pay:

  • Personal data. Names, email addresses, phone numbers, IBANs, national IDs and similar identifiers are removed before anything is licensed (how we protect your data).
  • DMs and private channels, excluded by default.
  • Data you process for your customers. If you’re a processor for your customers, as SaaS companies often are, only data you control as a company qualifies.
  • Client files in regulated professions. Accountants, lawyers and healthcare providers license only internal know-how (details).
  • Recent work. Nothing newer than 12 months, by default.

How selling works

  1. Get an estimate. Pick your data, team size and years. 60 seconds.
  2. Approve the scope. A 20-minute call. You pick the channels, folders and date ranges.
  3. We de-identify and list. Personal and confidential details are removed, and de-identified samples go into our catalog.
  4. Labs buy, you get paid. AI labs preview your samples and buy. You’re paid within 7 days of each sale.

That’s about an hour of your time in total. Your data stays under your control until a buyer purchases. Buyers get a non-exclusive license to a de-identified copy, and your originals stay yours. To compare grokked with other services, see the best data brokers for AI training data.

Questions

What kind of data do AI companies buy?

Data that shows real work and isn’t public: team chat and email, support tickets, documents and contracts, code and code reviews, CRM records, procedures, finance and project data. OpenAI asks for datasets “not already easily accessible online” (OpenAI).

How much is my company’s data worth to AI labs?

It depends on data types, volume, years of history and industry. Reported deals for one company’s workplace data range from $10,000 to several hundred thousand dollars (Gizmodo, Fortune). The grokked estimator gives a range for your company in 60 seconds.

Why would an AI lab pay for Slack messages and emails?

Conversations show how decisions get made and work gets done, which public web text rarely shows. Labs building AI agents need examples of people using tools, handling intermediate results and fixing mistakes.

Do I give up ownership of my data?

No. Buyers get a non-exclusive license to a de-identified copy, for training and evaluating AI models. Your originals never move and stay yours.

Is customer or employee data included?

No. Personal data is removed before anything is licensed, DMs and private channels are excluded, and data you process on behalf of your customers stays out.

See what your data is worth.

Get an estimate in 60 seconds. Listing your data costs nothing.

Get paid in 7 days

Sources

  1. OpenAI, November 9, 2023. OpenAI Data Partnerships
  2. Gizmodo, April 17, 2026. Failed companies are selling old Slack chats and email archives to train AI
  3. Fortune, September 14, 2026. Little-known AI startup Micro1 tries to trump Google’s bid for bankrupt Spirit Airlines’ data
  4. TechSpot, August 18, 2026. Google pays $10 million for 100 million Spirit Airlines emails and 500 million Teams chats to train AI
  5. Epoch AI, June 6, 2024. Will we run out of data to train large language models?
  6. The Verge, December 13, 2024. OpenAI cofounder Ilya Sutskever predicts the end of AI pre-training
  7. Anthropic. Non-user privacy policy
  8. Forbes, August 19, 2026. AI companies desperate for data are buying up dead airlines’ emails and scanning old books
  9. Reuters, February 22, 2024. Reddit in AI content licensing deal with Google, sources say
  10. TechCrunch, June 22, 2024. ‘What’s in it for us?’ journalists ask as publications sign content deals with AI firms
  11. Publishers Weekly, November 19, 2024. Agents, authors question HarperCollins AI deal
  12. TechCrunch, February 22, 2024. Reddit says it’s made $203M so far licensing its data
  13. Informa, May 8, 2024. Market update (RNS)
  14. Bloomberg Law, June 4, 2024. Shutterstock’s AI-licensing business generated $104 million
  15. Engadget, November 19, 2024. OpenAI will pay Dotdash Meredith at least $16 million per year to license its content
  16. TheWrap, July 30, 2025. New York Times seals $20 million AI deal with Amazon
  17. Engadget, March 3, 2026. Meta signs a multimillion dollar AI licensing deal with News Corp

About the author

Nico Vergauwen

Nico Vergauwen is the founder of grokked, a data broker that licenses companies’ de-identified internal data to AI labs. Owners never pay anything and are paid within 7 days of each sale.

More guides