A model team choosing between a USD 69,840 outside data engagement and a year of in-house labelling is making the core purchase of the AI training dataset market. The AI training dataset market covers licensed text, image, audio and sensor corpora plus custom collection, human annotation and reinforcement learning from human feedback (RLHF) sold to model builders. Douglas Insights sizes it at USD 3.27 billion in 2025, from 46,850 engagements at an average USD 69,840, and models USD 21.5 billion by 2035 at a 20.71% revenue CAGR. Disclosure now sits inside that price: the European Commission published its template for the public summary of training content on 24 July 2025, a common minimal baseline for general-purpose model providers. The study belongs to our enterprise software coverage and follows the Douglas Insights research methodology.
Why are frontier model labs and enterprises raising AI training dataset budgets so fast?
Engagement volume grows 15.4% a year to 2035 because post-training, enterprise fine-tuning and physical AI each need fresh, rights-cleared AI training datasets. Four drivers add 6.8, 4.3, 2.9 and 1.4 points, which sum to the volume leg. Price adds 4.6% a year on top as expert labelling displaces commodity tagging.
Post-training at frontier labs adds 6.8 points. Preference pairs, graded reasoning traces and red-team prompts are consumed in every model release cycle, and each cycle needs new AI training dataset batches rather than reused ones. Innodata reported full-year 2025 revenue of USD 251.7 million on 26 February 2026, 48% organic growth, and tied the demand to frontier model training, agentic AI and physical AI. A supplier growing at 48% in one year shows how hard the largest buyers are pulling.
Enterprise fine-tuning adds 4.3 points. Banks, insurers and retailers adapting open-weight models to their own documents buy smaller AI training dataset projects, typically domain annotation and evaluation sets. Douglas Insights counts these as 31.5% of 2025 value, or USD 1.03 billion, and the enterprise share of engagements rises because a single company now runs several model projects at once. Readers tracking the deployment side can compare our Artificial Intelligence Applications Market study.
Physical AI adds 2.9 points. Autonomous driving stacks, warehouse robots and humanoid programmes need lidar, radar and camera scenes labelled frame by frame, and those multimodal and sensor AI training datasets cost more per hour than text work. Autonomous vehicle developers account for 13.9% of 2025 value, about USD 454.8 million in our model.
Regulatory documentation adds the last 1.4 points. Under the EU AI Act, the obligations for general-purpose AI models became applicable on 2 August 2025, so providers need provenance records for every dataset they train on. Licensed corpora with clean chains of title gain share as a result, which is also why the 24 July 2025 template matters to procurement teams.
Add the four drivers together and engagements rise from 46,850 in 2025 to 54,065 in 2026. Demand is broad. No single buyer group supplies more than 6.8 of the 15.4 points, so one lab pausing its post-training budget slows the AI training dataset curve without breaking it. Price does the rest: expert raters cost more than generalists. Mix, not inflation, explains most of the 4.6% price leg.
Which copyright and quality headwinds slow AI training dataset spending?
Three headwinds remove 3.3 points from the 15.4% volume leg in our slower case: copyright exposure takes 1.6 points, synthetic substitution 1.1 points and annotator quality limits 0.6 points. None stops AI training dataset demand outright, but each lowers the 2035 value by billions of dollars.
Copyright exposure removes 1.6 points. Rights holders in the United States and Europe contest the use of scraped books, news and images, and buyers who cannot prove a licence face either settlement costs or retraining. Litigation is the bigger risk. Some labs respond by freezing purchases of web-derived AI training datasets until provenance is documented.
Synthetic substitution removes 1.1 points. Model-generated data is cheaper than human labels for code, mathematics and structured tasks, so part of the text AI training dataset budget moves into compute. Douglas Insights treats synthetic and tabular data as a segment, at 5.9% of 2025 value, because buyers still pay vendors to generate and verify it.
Annotator quality and supply remove 0.6 points. Experts are scarce. Expert RLHF work needs physicians, lawyers and engineers, and a shortage of qualified raters caps how many AI training dataset programmes a vendor can run in parallel.
Which data type leads AI training dataset revenue, text, imagery or speech?
Text leads with 41.3% of 2025 value, USD 1.35 billion, because every large language model release consumes new instruction, preference and evaluation sets. Image and video follows at 27.4%, then multimodal and sensor, audio and speech, and synthetic and tabular AI training datasets, which grow fastest at 27.4% a year.
| Data type | Share 2025 | Value 2025 | Growth 2026-2035 | Value 2035 |
|---|---|---|---|---|
| Text | 41.3% | USD 1.35 billion | 22.6% | USD 10.4 billion |
| Image and video | 27.4% | USD 896.5 million | 17.9% | USD 4.65 billion |
| Audio and speech | 11.8% | USD 386.1 million | 16.2% | USD 1.73 billion |
| Multimodal and sensor | 13.6% | USD 445.0 million | 19.1% | USD 2.56 billion |
| Synthetic and tabular | 5.9% | USD 193.1 million | 27.4% | USD 2.18 billion |
Text AI training datasets hold USD 1.35 billion and reach USD 10.4 billion by 2035 at 22.6% a year, since reasoning and agent training multiply the number of graded examples per model. Image and video work is worth USD 896.5 million and grows 17.9%, slower because bounding-box tagging is the most automated task in the field. Multimodal and sensor data, at USD 445.0 million, grows 19.1% on robotics and driving programmes. Audio and speech sets add USD 386.1 million at 16.2%, held back by mature transcription tools. Synthetic and tabular data is the smallest at USD 193.1 million, yet its 27.4% rate is the fastest because verification of generated data is billable work.
How do licensed corpora, custom collection and RLHF split AI training dataset spend?
Human annotation carries 34.6% of 2025 AI training dataset value, USD 1.13 billion, ahead of licensed corpora at 24.8%, reinforcement learning from human feedback at 22.3% and custom collection at 18.3%. The split shows buyers still pay mostly for human judgment rather than for raw content.
Licensed corpora, worth USD 811.5 million, are the purchase that legal teams prefer because the contract names the rights. RLHF programmes, at USD 729.7 million, carry the highest hourly rates in the AI training dataset market. Custom collection, at USD 598.8 million, covers speech recorded in target accents, store-shelf photography and driving scenes captured to order. By end user, frontier model developers buy 46.2% of value, USD 1.51 billion, enterprises 31.5%, autonomous vehicle developers 13.9% and healthcare 8.4%, or USD 274.8 million.
Where are AI training dataset contracts signed, and how fast does Asia Pacific grow?
North America signs the most AI training dataset value, USD 1.30 billion in 2025 or 39.7%, because most frontier labs buy there. Asia Pacific, at USD 975.1 million, grows fastest at 22.7% a year on Chinese model developers and Indian and Philippine annotation hubs.
| Region | 2025 | 2026 | 2035 | CAGR |
|---|---|---|---|---|
| Global | USD 3.27 billion | USD 3.95 billion | USD 21.5 billion | 20.71% |
| North America | USD 1.30 billion | USD 1.56 billion | USD 8.18 billion | 20.20% |
| Asia Pacific | USD 975.1 million | USD 1.20 billion | USD 7.54 billion | 22.70% |
| Europe | USD 664.2 million | USD 789.1 million | USD 3.72 billion | 18.80% |
| Latin America | USD 183.2 million | USD 222.8 million | USD 1.29 billion | 21.60% |
| Middle East and Africa | USD 150.5 million | USD 176.8 million | USD 753.1 million | 17.47% |
North America reaches USD 8.18 billion by 2035 at 20.2%, keeping first place but slipping to 38.1% of global value. Asia Pacific climbs to USD 7.54 billion, a 35.1% share, and Appen reported on 25 February 2026 that its China business earned USD 102.9 million in 2025, mostly from new and expanding LLM projects. Europe holds USD 664.2 million and grows 18.8%, with AI Act paperwork adding to project costs but also to licensed-data demand. Latin America, at USD 183.2 million, grows 21.6% as Spanish and Portuguese language data gains buyers. The Middle East and Africa is the wildcard: USD 150.5 million today, rising to USD 753.1 million at 17.47%, depending on how fast Gulf sovereign model programmes buy Arabic AI training datasets.
Which companies supply AI training datasets, and how concentrated is the competition?
The top three disclosed sellers, Innodata, Appen and Reddit, hold at most 19.0% of 2025 AI training dataset value on Douglas Insights estimates, treating each reported revenue line as an upper bound. Scale AI does not publish revenue, so true concentration is higher.
Scale AI announced on 12 June 2025 a Meta investment that values the company at over USD 29 billion, with founder Alexandr Wang joining Meta and Jason Droege becoming interim chief executive. Scale keeps its revenue private. Innodata earned USD 251.7 million in 2025 from data engineering for large model builders. Appen reported group operating revenue of USD 230.8 million for 2025, USD 127.9 million of it outside China.
Content owners sell AI training datasets too. Reddit reported on 5 February 2026 full-year 2025 revenue of USD 2.20 billion, of which other revenue, including data licensing, was USD 140 million. Shutterstock reported 2024 data revenue of about USD 120.3 million, up 15%, inside its data, distribution and services line. TELUS Digital, iMerit, Labelbox and Snorkel AI complete the field of annotation workforces and labelling software vendors.
| Company | Role | Disclosed figure |
|---|---|---|
| Innodata | Data engineering for model builders | USD 251.7 million revenue, 2025 |
| Appen | Annotation and LLM projects | USD 230.8 million revenue, 2025 |
| Licensor of conversation data | USD 140 million other revenue, 2025 | |
| Shutterstock | Licensor of image and video data | USD 120.3 million data revenue, 2024 |
| Scale AI | RLHF and autonomy labelling | Valued at over USD 29 billion, 2025 |
How much does an AI training dataset engagement cost, from bounding boxes to expert RLHF?
The average AI training dataset engagement costs USD 69,840 in 2025 and rises 4.6% a year to USD 109,502 by 2035, because expert reasoning labels replace commodity image tags in the mix. Bands run from a few thousand dollars for small tagging jobs to seven figures for RLHF programmes.
Douglas Insights puts realised bands at USD 8,000 to 25,000 for commodity bounding-box and transcription contracts, USD 40,000 to 250,000 for domain annotation and evaluation sets, and USD 150,000 to 1.2 million for multi-month expert RLHF programmes. Corpus licences sit outside those bands: a single content owner such as Reddit reports USD 140 million a year of other revenue. For 2026 the model uses USD 73,053 per engagement. Quotes vary widely.
Buyers comparing quotes should price the hidden items: rework rates, inter-annotator agreement guarantees and provenance documentation for the AI Act summary. On a USD 69,840 engagement those items often decide whether outsourcing beats an internal team.
Which AI Act and copyright rules must AI training dataset providers comply with?
Since 2 August 2025, general-purpose model providers in the EU must meet AI Act obligations, including a public summary of training content, so every AI training dataset purchase now needs documented sources. The AI Act entered into force on 1 August 2024.
The Commission template asks providers to describe data sources by category, which pushes AI training dataset buyers toward vendors that supply licence records, collection dates and opt-out handling. In the United States, copyright claims over scraped works are being tested in court rather than by statute, so contracts carry indemnities. Personal data inside speech, medical and driving datasets also falls under privacy law in each of the 5 regions we model.
What if synthetic data or new licensing deals reshape the AI training dataset forecast to 2035?
The base case reaches USD 21.5 billion in 2035, with 15.4% volume growth and 4.6% price growth; the slower case gives USD 14.3 billion and the faster case USD 29.8 billion. Synthetic substitution and copyright rulings decide which AI training dataset path holds.
The slower case runs 12.1% volume and 3.4% price, the outcome if courts force licence fees onto scraped data and labs answer with synthetic data. The faster case runs 18.2% volume and 5.5% price, if physical AI programmes scale and the EU disclosure rules turn licensed data into the default purchase. A single extra point of engagement growth adds about USD 1.94 billion to the 2035 AI training dataset value.
Published forecasts run from about 20.3% to 28.9% a year. Our 20.71% sits near the lower end because we count only external spend, not in-house labelling.
Douglas Exclusive: the AI Training Dataset Sourcing Risk Matrix
The AI Training Dataset Sourcing Risk Matrix is a Douglas Insights model built from 12 inputs: 7 primary documents read for this study and 5 published growth forecasts. It scores each data type against each sourcing route for legal exposure on a 1 to 5 scale, where 5 is the highest risk.
| Data type | Licensed corpora | Custom collection | Human annotation | RLHF |
|---|---|---|---|---|
| Text | 2 | 2 | 1 | 1 |
| Image and video | 3 | 2 | 1 | 2 |
| Audio and speech | 3 | 4 | 2 | 2 |
| Multimodal and sensor | 2 | 3 | 1 | 2 |
| Synthetic and tabular | 1 | 1 | 1 | 2 |
The matrix finds that 6 of its 20 cells score 3 or more, and those cells hold the AI training dataset purchases most exposed to the EU summary template and to copyright claims. Consent is the weak point. Speech recorded to order scores 4 because it combines voice biometrics with consent records. Annotation of buyer-owned data scores lowest, which explains why human annotation keeps a 34.6% share despite cheaper synthetic options.
Which buyers should build AI training datasets in-house rather than buy?
Teams spending more than about USD 1.2 million a year on one AI training dataset type, the top of the RLHF band, often gain by hiring raters directly. Volume decides. Below that level, outside vendors win on recruitment speed and quality tooling.
Douglas Insights expects external engagements to reach 196,231 by 2035, from 46,850 in 2025, because most enterprises sit below that threshold. Infrastructure buyers can read the related Data Centre Liquid Cooling Market report, and data teams the Enterprise Data Observability Platforms Market study.
How the model was built: which inputs back the AI training dataset receipts?
The AI training dataset model multiplies 46,850 engagements by USD 69,840 to give USD 3.27 billion for 2025, across 5 data types, 4 sourcing routes, 4 end users and 5 regions. Volume grows 15.4% and price 4.6% a year, which multiply to the 20.71% revenue CAGR.
Three cross-checks support the base. First, published 2025 estimates near USD 3.2 billion sit 2.25% below our figure. Second, the four disclosed seller revenues above total about USD 742.8 million, roughly 22.7% of our 2025 value, a plausible share for four firms in a fragmented field. Third, our 20.71% CAGR falls inside the 20.3% to 28.9% published range. Regions sum to the global value in both 2025 and 2035, and the five data-type values sum to USD 3.27 billion.
How this report is built
- Every figure carries a confidence grade in the fact sheet above, and the working model ships with every licence.
- Five regional models sum to the global figure, with country tables in the Excel model.
- The next scheduled review of this study is April 2027.
- Licence holders receive it as a maintained tab in the Excel model.
Sources
- European Commission Explanatory notice and template for the public summary of training content (2025)
- European Commission AI Act regulatory framework (2025)
- Appen Limited (ASX) FY25 results and FY26 outlook (2026)
- Innodata Inc. Fourth quarter and full year 2025 results (2026)
- Reddit, Inc. Q4 2025 earnings press release (2026)
- Scale AI Scale AI announces next phase of company evolution (2025)
- Shutterstock, Inc. Full year 2024 and fourth quarter financial results (2025)
Inside the 196-page report
01Executive summary12 sections
The market in one view
- 1.1Market snapshot, 2025 and 2035
- 1.1.1Market size, 2025
- 1.1.2Forecast, 2035
- 1.1.3Growth rate, 2026–2035
- 1.2Growth decomposition
- 1.2.1Volume growth (thousand engagements)
- 1.2.2Value per unit growth
- 1.3Key findings
- 1.4Segment highlights
- 1.5Regional highlights
- 1.6Competitive highlights
- 1.7Douglas Insights verdict
02Scope and definitions17 sections
What the study counts
- 2.1Market definition
- 2.2Inclusions and exclusions
- 2.2.1Data types
- 2.2.2Sourcing routes
- 2.2.3Exclusions
- 2.3Segmentation
- 2.3.1By data type
- 2.3.2By sourcing model
- 2.3.3By end user
- 2.3.4By region
- 2.4Years considered
- 2.4.1Base year 2025
- 2.4.2Forecast 2026–2035
- 2.5Currency and units
- 2.5.1Value in USD million
- 2.5.2Volume in thousand engagements
- 2.6Who this report is for
03Research methodology16 sections
Bottom-up: thousand engagements × value per unit
- 3.1Bottom-up market model
- 3.1.1Volume base, 2025 (thousand engagements)
- 3.1.2Value per unit
- 3.1.3Forecast legs to 2035
- 3.2Top-down cross-checks
- 3.3Data triangulation
- 3.4Sources
- 3.4.1Regulators and statistics offices
- 3.4.2Company filings and results
- 3.4.3Trade and industry bodies
- 3.4.47 primary sources cited
- 3.5Confidence grading
- 3.6Assumptions and limitations
- 3.6.1Unit build
- 3.6.2Price build
- 3.6.3Cross-checks
04Growth drivers3 sections
Post-training, enterprise fine-tuning, physical AI
- 4.1Driver contributions
- 4.2Frontier lab demand
- 4.3Regulatory pull
05Restraints3 sections
Copyright, synthetic data, rater supply
- 5.1Points removed
- 5.2Litigation exposure
- 5.3Quality limits
06Segments by data type3 sections
Text, image and video, audio, multimodal, synthetic
- 6.1Shares
- 6.2Growth rates
- 6.32035 values
07End users3 sections
Frontier labs, enterprises, autonomy, healthcare
- 7.1Shares
- 7.2Purchase patterns
- 7.3Outlook
08Pricing3 sections
Engagement value bands
- 8.1Commodity tagging
- 8.2Domain annotation
- 8.3Expert RLHF
09Regulation3 sections
AI Act and copyright
- 9.1EU summary template
- 9.2US litigation
- 9.3Privacy
10Build or buy guidance3 sections
When in-house labelling wins
- 10.1Spend threshold
- 10.2Vendor selection
- 10.3Contract terms
11Market size and forecast, 2025–20355 sections
Global value, volume and value per unit
- 11.1Market value, 2025–2035
- 11.2Volume (thousand engagements), 2025–2035
- 11.3Value per unit, 2025–2035
- 11.4Year-on-year growth
- 11.5Growth decomposition
12AI Training Dataset market, by data type16 sections
5 segments, value 2025–2035
- 12.1Overview and share, 2025 and 2035
- 12.2Text
- 12.2.1Market size and forecast, 2025–2035
- 12.2.2Growth outlook
- 12.3Image and video
- 12.3.1Market size and forecast, 2025–2035
- 12.3.2Growth outlook
- 12.4Audio and speech
- 12.4.1Market size and forecast, 2025–2035
- 12.4.2Growth outlook
- 12.5Multimodal and sensor
- 12.5.1Market size and forecast, 2025–2035
- 12.5.2Growth outlook
- 12.6Synthetic and tabular
- 12.6.1Market size and forecast, 2025–2035
- 12.6.2Growth outlook
13AI Training Dataset market, by sourcing model13 sections
4 segments, value 2025–2035
- 13.1Overview and share, 2025 and 2035
- 13.2Licensed corpora
- 13.2.1Market size and forecast, 2025–2035
- 13.2.2Growth outlook
- 13.3Custom collection
- 13.3.1Market size and forecast, 2025–2035
- 13.3.2Growth outlook
- 13.4Human annotation
- 13.4.1Market size and forecast, 2025–2035
- 13.4.2Growth outlook
- 13.5Reinforcement learning from human feedback
- 13.5.1Market size and forecast, 2025–2035
- 13.5.2Growth outlook
14AI Training Dataset market, by end user13 sections
4 segments, value 2025–2035
- 14.1Overview and share, 2025 and 2035
- 14.2Frontier model developers
- 14.2.1Market size and forecast, 2025–2035
- 14.2.2Growth outlook
- 14.3Enterprises
- 14.3.1Market size and forecast, 2025–2035
- 14.3.2Growth outlook
- 14.4Autonomous vehicle developers
- 14.4.1Market size and forecast, 2025–2035
- 14.4.2Growth outlook
- 14.5Healthcare
- 14.5.1Market size and forecast, 2025–2035
- 14.5.2Growth outlook
15Regional analysis26 sections
5 regions
- 15.1Regional overview and share, 2025 and 2035
- 15.2North America
- 15.2.1Market size and forecast, 2025–2035
- 15.2.2By data type
- 15.2.3By sourcing model
- 15.2.4By end user
- 15.3Asia Pacific
- 15.3.1Market size and forecast, 2025–2035
- 15.3.2By data type
- 15.3.3By sourcing model
- 15.3.4By end user
- 15.4Europe
- 15.4.1Market size and forecast, 2025–2035
- 15.4.2By data type
- 15.4.3By sourcing model
- 15.4.4By end user
- 15.5Latin America
- 15.5.1Market size and forecast, 2025–2035
- 15.5.2By data type
- 15.5.3By sourcing model
- 15.5.4By end user
- 15.6Middle East and Africa
- 15.6.1Market size and forecast, 2025–2035
- 15.6.2By data type
- 15.6.3By sourcing model
- 15.6.4By end user
16Competitive landscape9 sections
5 companies profiled
- 16.1Market concentration
- 16.2Market share analysis, 2025
- 16.3Strategic moves: acquisitions, launches, contracts
- 16.4Company profilesEach profile: overview, products, financials where reported, position in this market, recent developments
- 16.4.1Innodata
- 16.4.2Appen
- 16.4.3Reddit
- 16.4.4Shutterstock
- 16.4.5Scale AI
17Scenarios to 20355 sections
Slower, base and faster cases
- 17.1Slower case
- 17.2Base case case
- 17.3Faster case
- 17.4Sensitivity of the 2035 value
- 17.5Published forecasts compared
18Douglas Exclusive: the AI Training Dataset Sourcing Risk Matrix3 sections
Legal exposure by data type and route
- 18.1Scoring
- 18.2Findings
- 18.3Inputs
19Appendix5 sections
Data, sources and licence
- 19.1Data tables (Excel model)
- 19.2Sources (7)
- 19.3Abbreviations
- 19.4Change log and next review
- 19.5Licence and how to cite
TList of tables38
- Table 1Market value, 2025–2035 (USD million)
- Table 2Volume, 2025–2035 (thousand engagements)
- Table 3Value per unit, 2025–2035
- Table 4AI Training Dataset market by data type, 2025–2035 (USD million)
- Table 5Text: market size, 2025–2035 (USD million)
- Table 6Image and video: market size, 2025–2035 (USD million)
- Table 7Audio and speech: market size, 2025–2035 (USD million)
- Table 8Multimodal and sensor: market size, 2025–2035 (USD million)
- Table 9Synthetic and tabular: market size, 2025–2035 (USD million)
- Table 10AI Training Dataset market by sourcing model, 2025–2035 (USD million)
- Table 11Licensed corpora: market size, 2025–2035 (USD million)
- Table 12Custom collection: market size, 2025–2035 (USD million)
- Table 13Human annotation: market size, 2025–2035 (USD million)
- Table 14Reinforcement learning from human feedback: market size, 2025–2035 (USD million)
- Table 15AI Training Dataset market by end user, 2025–2035 (USD million)
- Table 16Frontier model developers: market size, 2025–2035 (USD million)
- Table 17Enterprises: market size, 2025–2035 (USD million)
- Table 18Autonomous vehicle developers: market size, 2025–2035 (USD million)
- Table 19Healthcare: market size, 2025–2035 (USD million)
- Table 20AI Training Dataset market by region, 2025–2035 (USD million)
- Table 21North America: market by data type, 2025–2035 (USD million)
- Table 22North America: market by sourcing model, 2025–2035 (USD million)
- Table 23North America: market by end user, 2025–2035 (USD million)
- Table 24Asia Pacific: market by data type, 2025–2035 (USD million)
- Table 25Asia Pacific: market by sourcing model, 2025–2035 (USD million)
- Table 26Asia Pacific: market by end user, 2025–2035 (USD million)
- Table 27Europe: market by data type, 2025–2035 (USD million)
- Table 28Europe: market by sourcing model, 2025–2035 (USD million)
- Table 29Europe: market by end user, 2025–2035 (USD million)
- Table 30Latin America: market by data type, 2025–2035 (USD million)
- Table 31Latin America: market by sourcing model, 2025–2035 (USD million)
- Table 32Latin America: market by end user, 2025–2035 (USD million)
- Table 33Middle East and Africa: market by data type, 2025–2035 (USD million)
- Table 34Middle East and Africa: market by sourcing model, 2025–2035 (USD million)
- Table 35Middle East and Africa: market by end user, 2025–2035 (USD million)
- Table 36Company market shares, 2025
- Table 37Scenario values, 2035
- Table 38Sources and confidence grades by figure
FList of figures9
- Figure 1Market value, 2025–2035
- Figure 2Growth decomposition, 2026–2035
- Figure 3Share by data type, 2025 and 2035
- Figure 4Share by sourcing model, 2025 and 2035
- Figure 5Share by end user, 2025 and 2035
- Figure 6Share by region, 2025 and 2035
- Figure 7Growth by region, 2026–2035
- Figure 8Market concentration, 2025
- Figure 9Scenario paths to 2035
Questions buyers ask
What does an average outsourced AI training data engagement cost?
USD 69,840 in 2025 on Douglas Insights estimates, rising to USD 109,502 by 2035 as expert RLHF work replaces commodity image tagging in the purchase mix.
How much revenue do AI training datasets generate today and in 2035?
USD 3.27 billion in 2025, from 46,850 engagements, and USD 21.5 billion by 2035 in the base case, a 20.71% revenue CAGR.
Why does text data take the biggest slice of dataset spending?
41.3% of 2025 value, USD 1.35 billion, goes to text because every language model release needs new instruction, preference and evaluation sets.
Is synthetic data replacing human-labelled training data?
5.9% of 2025 value is synthetic and tabular data today, growing 27.4% a year, while synthetic substitution removes 1.1 points of volume growth in our slower case.
What did the EU publish for training data disclosure?
24 July 2025: the European Commission published its template for the public summary of training content, after AI Act obligations for general-purpose models became applicable on 2 August 2025.
How concentrated is supply of training data?
19.0% at most is held by the top three disclosed sellers, Innodata, Appen and Reddit, on Douglas Insights estimates; Scale AI does not disclose revenue.
Which region adds training data spend fastest?
22.7% a year in Asia Pacific, from USD 975.1 million in 2025, driven by Chinese model developers; Appen China earned USD 102.9 million in 2025.
What separates the slower and faster outcomes for 2035?
USD 14.3 billion in the slower case versus USD 29.8 billion in the faster case, depending mainly on copyright rulings and synthetic substitution.
Research & citation
This report was researched, written and reviewed by the Douglas Insights Research Desk under the Douglas Insights editorial standards. Material errors are logged in the corrections log. No section is sponsored.
Douglas Insights Inc (2026). AI Training Dataset Market. Report DI-IT-10602, October 2026. https://www.douglasinsights.com/ai-training-dataset-market/