What company data AI labs pay for, and what it’s worth
By Nico Vergauwen, founder of grokkedPublished Updated 6 min read
AI labs pay for data that shows how work actually gets done: team chat and email threads, support tickets, documents and contracts, code reviews, CRM pipelines, procedures and finance workflows. That data isn’t on the public web, which labs have largely used up. OpenAI says it wants datasets “not already easily accessible online to the public today” (OpenAI).
The prices are real but not public. Reported deals for one company’s workplace data run from $10,000 to several hundred thousand dollars (Gizmodo, Fortune), and Google won a bankruptcy auction for Spirit Airlines’ emails, Teams chats and files at $10 million in 2026. Below: the data types that sell, why private data is worth more, what labs have paid and how to estimate your own.
Key takeaways
- The most valuable company data shows people solving real problems step by step: support tickets, code reviews and incidents, procedures, and the chat and email around them.
- Public text is running out. Epoch AI puts the stock at about 300 trillion tokens and expects models to fully use it between 2026 and 2032.
- Labs have paid about $60 million a year for Reddit’s data, a reported $250 million+ over five years for News Corp’s, and $5,000 per book in a HarperCollins deal.
- What your data earns depends on data types, team size, years of history and industry. The grokked estimator gives a range in 60 seconds.
The data types AI labs buy
Labs train models to do professional work, so they want records of professional work. These are the eight types grokked licenses, with the tools they usually live in:
| Data | Examples | Usually in | What a model learns from it |
|---|---|---|---|
| Customer support | Tickets, chats, help-desk replies | Zendesk, Freshdesk, Intercom | How real questions get diagnosed, answered and escalated |
| Sales & CRM | Pipelines, call notes, proposals | Salesforce, HubSpot, Pipedrive | How deals are qualified, priced and negotiated |
| Code & engineering | Repositories, code reviews, incidents | GitHub, GitLab, Jira | How software is written, reviewed and fixed |
| Procedures & SOPs | Wikis, manuals, checklists | Confluence, Notion, SharePoint | Step-by-step methods for a trade |
| Contracts & documents | Contracts, reports, bids | Google Drive, Word, DocuSign | How professional documents are structured and argued |
| Chat & email | Team channels, email threads | Slack, Teams, Outlook | How decisions get made in conversation |
| Finance & accounting | Bookkeeping, invoices, orders | Exact, SAP, Odoo | How transactions are recorded, checked and closed |
| Projects & tasks | Plans, tasks, status updates | Asana, monday.com, ClickUp | How work is planned and tracked over time |
The Spirit Airlines auction shows the appetite. The dataset holds about 100 million emails, 500 million Teams chats, 17 million OneDrive files, more than 20 million SharePoint items and 516 code repositories, with personal information removed before transfer (TechSpot).
Workflow data matters most for AI agents. As Fortune put it, “an agent needs to learn how to choose actions, use tools, respond to intermediate results and recover from mistakes” (Fortune). A support ticket that goes from complaint to fix, or a code review that catches a bug, shows exactly that.
Why private data is worth more than public data
Most of the public web has already gone into training. Epoch AI estimates the stock of quality public human text at around 300 trillion tokens and expects models to fully use it between 2026 and 2032 (Epoch AI). “We’ve achieved peak data and there’ll be no more,” OpenAI co-founder Ilya Sutskever said at NeurIPS in December 2024. “There’s only one internet” (The Verge).
So labs look for data nobody else has. OpenAI’s data partnerships page says it is “particularly looking for data that expresses human intention (e.g., long-form writing or conversations rather than disconnected snippets)” (OpenAI). Anthropic’s privacy policy lists “non-public datasets obtained from businesses” among its training sources (Anthropic).
Forbes summed up the shift in August 2026: “Enterprise data has become incredibly useful for AI training, since labs had already hoovered up the entire public internet by late 2024” (Forbes).
What AI labs have paid for data
Those are the headline deals. Other disclosed prices:
| Seller | Buyer | Reported price | Source |
|---|---|---|---|
| AI companies | $203M in data licensing contracts over two to three years | TechCrunch, from Reddit’s IPO filing | |
| Taylor & Francis (Informa) | Microsoft | $10M+ access fee, plus recurring payments in 2025–2027 | Informa |
| Shutterstock | AI companies | $104M in AI licensing revenue in 2023 | Bloomberg Law |
| Dotdash Meredith | OpenAI | At least $16M a year | Engadget |
| The New York Times | Amazon | $20M–$25M a year | TheWrap |
| News Corp | Meta | Up to $50M a year, for three years | Engadget |
| Spirit Airlines, in bankruptcy | $10M at auction; Micro1 later offered $12.5M | Fortune | |
| Startups closing down | AI labs and data firms | $10,000–$100,000 per company | Gizmodo |
Publisher deals pay for public-facing content. Company data is a newer market: per-company prices are reported, not listed, and “there is no reliable market average,” Fortune noted (Fortune).
What your company’s data could be worth
grokked’s estimator values each data type separately, then adjusts for team size, years of history and industry. Value grows with headcount, but less than proportionally, and we count the most recent 15 years. Data that shows step-by-step work, such as code, support tickets and procedures, carries the most training signal per person.
| Example company | Data | Estimated payout |
|---|---|---|
| 20-person agency, 8 years | Chat & email, Contracts & documents, Projects & tasks | €30K–€43K |
| 40-person accounting firm, 20 years | Procedures & SOPs, Contracts & documents, Chat & email | €72K–€105K |
| 60-person engineering firm, 15 years | Chat & email, Contracts & documents, Procedures & SOPs, Projects & tasks | €155K–€225K |
| 150-person software company, 10 years | Code & engineering, Customer support, Chat & email, Procedures & SOPs | €295K–€435K |
| 500-person logistics company, 12 years | Customer support, Sales & CRM, Procedures & SOPs, Finance & accounting | €605K–€885K |
Indicative. The final amount depends on volume, quality and what buyers pay. For accountants, lawyers and healthcare, client files are excluded, and the estimate already accounts for that.
You’re paid per sale. AI labs preview de-identified samples in our catalog and license your dataset, and you’re paid within 7 days of each sale. You never pay anything; grokked takes a commission on each sale.
See what your data is worth.
Get an estimate in 60 seconds. Listing your data costs nothing.
What stays out
Some data never sells, whatever a buyer would pay:
- Personal data. Names, email addresses, phone numbers, IBANs, national IDs and similar identifiers are removed before anything is licensed (how we protect your data).
- DMs and private channels, excluded by default.
- Data you process for your customers. If you’re a processor for your customers, as SaaS companies often are, only data you control as a company qualifies.
- Client files in regulated professions. Accountants, lawyers and healthcare providers license only internal know-how (details).
- Recent work. Nothing newer than 12 months, by default.
How selling works
- Get an estimate. Pick your data, team size and years. 60 seconds.
- Approve the scope. A 20-minute call. You pick the channels, folders and date ranges.
- We de-identify and list. Personal and confidential details are removed, and de-identified samples go into our catalog.
- Labs buy, you get paid. AI labs preview your samples and buy. You’re paid within 7 days of each sale.
That’s about an hour of your time in total. Your data stays under your control until a buyer purchases. Buyers get a non-exclusive license to a de-identified copy, and your originals stay yours. To compare grokked with other services, see the best data brokers for AI training data.
Questions
What kind of data do AI companies buy?
Data that shows real work and isn’t public: team chat and email, support tickets, documents and contracts, code and code reviews, CRM records, procedures, finance and project data. OpenAI asks for datasets “not already easily accessible online” (OpenAI).
How much is my company’s data worth to AI labs?
It depends on data types, volume, years of history and industry. Reported deals for one company’s workplace data range from $10,000 to several hundred thousand dollars (Gizmodo, Fortune). The grokked estimator gives a range for your company in 60 seconds.
Why would an AI lab pay for Slack messages and emails?
Conversations show how decisions get made and work gets done, which public web text rarely shows. Labs building AI agents need examples of people using tools, handling intermediate results and fixing mistakes.
Do I give up ownership of my data?
No. Buyers get a non-exclusive license to a de-identified copy, for training and evaluating AI models. Your originals never move and stay yours.
Is customer or employee data included?
No. Personal data is removed before anything is licensed, DMs and private channels are excluded, and data you process on behalf of your customers stays out.
See what your data is worth.
Get an estimate in 60 seconds. Listing your data costs nothing.
Sources
- OpenAI, November 9, 2023. OpenAI Data Partnerships
- Gizmodo, April 17, 2026. Failed companies are selling old Slack chats and email archives to train AI
- Fortune, September 14, 2026. Little-known AI startup Micro1 tries to trump Google’s bid for bankrupt Spirit Airlines’ data
- TechSpot, August 18, 2026. Google pays $10 million for 100 million Spirit Airlines emails and 500 million Teams chats to train AI
- Epoch AI, June 6, 2024. Will we run out of data to train large language models?
- The Verge, December 13, 2024. OpenAI cofounder Ilya Sutskever predicts the end of AI pre-training
- Anthropic. Non-user privacy policy
- Forbes, August 19, 2026. AI companies desperate for data are buying up dead airlines’ emails and scanning old books
- Reuters, February 22, 2024. Reddit in AI content licensing deal with Google, sources say
- TechCrunch, June 22, 2024. ‘What’s in it for us?’ journalists ask as publications sign content deals with AI firms
- Publishers Weekly, November 19, 2024. Agents, authors question HarperCollins AI deal
- TechCrunch, February 22, 2024. Reddit says it’s made $203M so far licensing its data
- Informa, May 8, 2024. Market update (RNS)
- Bloomberg Law, June 4, 2024. Shutterstock’s AI-licensing business generated $104 million
- Engadget, November 19, 2024. OpenAI will pay Dotdash Meredith at least $16 million per year to license its content
- TheWrap, July 30, 2025. New York Times seals $20 million AI deal with Amazon
- Engadget, March 3, 2026. Meta signs a multimillion dollar AI licensing deal with News Corp
About the author
Nico Vergauwen
Nico Vergauwen is the founder of grokked, a data broker that licenses companies’ de-identified internal data to AI labs. Owners never pay anything and are paid within 7 days of each sale.