Machine learning foundations
The methods behind learning: generalization, representations, optimization, causality, uncertainty and data.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Learning theory
Generalization, sample complexity, statistical learning, and the theory of deep and in-context learning.
- Authors · observed
- 1,378
- Investment · scenario
- $107M–$430M
Sources and methodology
Publication activity
- 2019
- 185
- 2020
- 281
- 2021
- 324
- 2022
- 340
- 2023
- 397
- 2024
- 448
- 2025
- 575
The query is restricted to AI and ML; it does not cover all statistical learning theory or mathematical theory.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("statistical learning theory" OR "generalization bound" OR "sample complexity" OR "neural tangent kernel" OR "grokking" OR "theory of deep learning" OR "PAC learning") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
1,378 distinct OpenAlex Author IDs across all 575 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 575 of 575 papers (100.0%). 14.7% of authorship records lack an Author ID; 13.6% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $215M.
- AI infrastructure/models/research/governance: $143.2B × 0.15%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers learning guarantees. Training algorithms are listed under optimization.
Organizations
Selected examples- Google Research Research
Learning theory, optimization, reinforcement learning and differential privacy research.
- Microsoft Research
Microsoft Research studies theoretical foundations of machine learning.
- Amazon / AWS Research
Research establishes generalization guarantees for learned ensemble strategies.
Training optimization
Stochastic and distributed optimizers, learning-rate schedules, training stability, and matrix methods.
- Authors · scenario
- 3.2k–6.7k
- Investment · scenario
- $430M–$1.72B
Sources and methodology
Publication activity
- 2019
- 581
- 2020
- 799
- 2021
- 995
- 2022
- 1,036
- 2023
- 1,166
- 2024
- 1,460
- 2025
- 2,045
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("stochastic gradient descent" OR "neural network optimization" OR "adaptive optimizer" OR "learning rate schedule" OR "sharpness aware minimization" OR "second order optimization") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.2k publishing authors. Paper count × 3.29 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,188 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,045 papers (48.9%). 13.2% of authorship records lack an Author ID; 20.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.29; 95% bootstrap interval 3.13–3.44. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $859M.
- AI infrastructure/models/research/governance: $143.2B × 0.60%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on learning model parameters. Scheduling and routing problems sit under combinatorial optimization.
Organizations
Selected examples- Google Research Research
Learning theory, optimization, reinforcement learning and differential privacy research.
- Microsoft Research
Research on optimization methods for machine learning.
- Google DeepMind Research
Optax provides gradient transformations and optimization algorithms.
More organizations (1)
- Meta Research
Schedule-free optimization studies training methods that do not require learning-rate decay schedules.
Representation learning
Self-supervised, contrastive and semi-supervised learning, clustering, and dimensionality reduction.
- Authors · scenario
- 15k–73k
- Investment · scenario
- $2.11B–$8.46B
Sources and methodology
Publication activity
- 2019
- 3,062
- 2020
- 4,729
- 2021
- 6,676
- 2022
- 8,888
- 2023
- 11,349
- 2024
- 14,113
- 2025
- 16,868
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("self supervised learning" OR "representation learning" OR "contrastive learning" OR "semi supervised learning" OR "deep clustering" OR "unsupervised learning")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 24k publishing authors. Paper count × 4.30 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,227 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 16,868 papers (5.9%). 11.8% of authorship records lack an Author ID; 26.0% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.30; 95% bootstrap interval 4.14–4.46. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $4.23B.
- AI infrastructure/models/research/governance: $143.2B × 1.00%
- Pharmaceutical: $10.6B × 15.00%
- Biotech: $4.84B × 25.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Methods used across data types; application areas provide a separate view.
Organizations
Selected examples- Meta Research
DINOv3 provides self-supervised visual representations.
- OpenAI Research
CLIP learns visual representations through image-text contrastive training.
- Google Research Research
SimCLR research learns visual representations through contrastive training.
More organizations (1)
- Apple Machine Learning Research Research
MobileCLIP studies efficient contrastive image-text representation learning.
Transfer and continual learning
Domain adaptation, few-shot and meta-learning, continual learning, and distribution shift.
- Authors · scenario
- 16k–81k
- Investment · scenario
- $358M–$1.43B
Sources and methodology
Publication activity
- 2019
- 3,576
- 2020
- 6,137
- 2021
- 8,757
- 2022
- 10,794
- 2023
- 13,205
- 2024
- 16,094
- 2025
- 20,365
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("transfer learning" OR "domain adaptation" OR "continual learning" OR "meta learning" OR "few shot learning" OR "domain generalization")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 27k publishing authors. Paper count × 3.97 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,928 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 20,365 papers (4.9%). 11.3% of authorship records lack an Author ID; 21.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.97; 95% bootstrap interval 3.76–4.23. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $716M.
- AI infrastructure/models/research/governance: $143.2B × 0.50%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers transferring and retaining knowledge. Model-specific fine-tuning also appears under post-training.
Organizations
Selected examples- OpenAI Research
CLIP research evaluates zero-shot transfer from image-text pretraining.
- Google Research Research
Meta-Dataset studies few-shot learning and generalization to unseen datasets.
- Physical Intelligence Research
Generalist robot policies learned across tasks and robot embodiments.
More organizations (1)
- Apple Machine Learning Research Research
MobileCLIP evaluates zero-shot classification and retrieval across downstream datasets.
Probabilistic ML and uncertainty
Bayesian inference, graphical models, Monte Carlo, probabilistic programming, calibration, and conformal prediction.
- Authors · scenario
- 3.9k–19k
- Investment · scenario
- $1.56B–$6.24B
Sources and methodology
Publication activity
- 2019
- 824
- 2020
- 1,303
- 2021
- 1,713
- 2022
- 2,181
- 2023
- 2,549
- 2024
- 3,202
- 2025
- 4,617
Many matches apply probabilistic methods. This count does not measure the size of the methods research community.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("Bayesian neural network" OR "Gaussian process" OR "probabilistic graphical model" OR "probabilistic programming" OR "uncertainty quantification" OR "conformal prediction") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.3k publishing authors. Paper count × 4.11 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,926 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,617 papers (21.7%). 9.2% of authorship records lack an Author ID; 24.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.11; 95% bootstrap interval 3.89–4.38. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.12B.
- AI infrastructure/models/research/governance: $143.2B × 0.20%
- Medical and healthcare: $11.8B × 5.00%
- Pharmaceutical: $10.6B × 10.00%
- Biotech: $4.84B × 15.00%
- Energy management: $4.64B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Causal inference is listed separately. Uncertainty research extends beyond language-model hallucinations.
Organizations
Selected examples- PyMC Labs Tools & infrastructure
PyMC Labs develops Bayesian modeling tools and supports probabilistic inference with PyMC.
- Amazon / AWS Research
Published research on probabilistic forecasting and time series.
- Microsoft Tools & infrastructure
Infer.NET provides probabilistic programming and Bayesian inference.
More organizations (1)
- Google Research Research
Research develops automated structured variational inference for probabilistic programs.
Causal inference
Causal discovery, causal representations, intervention effects, and uplift modeling.
- Authors · scenario
- 3.5k–6.3k
- Investment · scenario
- $1.32B–$5.28B
Sources and methodology
Publication activity
- 2019
- 186
- 2020
- 286
- 2021
- 391
- 2022
- 505
- 2023
- 588
- 2024
- 849
- 2025
- 1,661
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("causal inference" OR "causal discovery" OR "causal representation" OR "causal machine learning" OR "treatment effect" OR "uplift modeling") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.5k publishing authors. Paper count × 3.79 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,470 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,661 papers (60.2%). 12.3% of authorship records lack an Author ID; 19.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.79; 95% bootstrap interval 3.53–4.05. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.64B.
- AI infrastructure/models/research/governance: $143.2B × 0.15%
- Marketing, digital ads: $2.51B × 5.00%
- Medical and healthcare: $11.8B × 5.00%
- Pharmaceutical: $10.6B × 10.00%
- Fintech: $6.52B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on cause and effect rather than prediction alone.
Organizations
Selected examples- Microsoft Research
Causal machine learning.
- Uber Tools & infrastructure
Uber-origin CausalML provides uplift modeling and treatment-effect estimation.
- PyMC Labs Tools & infrastructure
CausalPy supports Bayesian causal analysis of quasi-experiments.
More organizations (1)
- IBM Tools & infrastructure
Causal 360 provides methods for causal inference and intervention analysis.
AutoML and experimental design
Model selection, hyperparameter search, Bayesian optimization, active learning, and choosing the next experiment.
- Authors · scenario
- 4.1k–20k
- Investment · scenario
- $1.79B–$7.17B
Sources and methodology
Publication activity
- 2019
- 751
- 2020
- 1,256
- 2021
- 1,780
- 2022
- 2,026
- 2023
- 2,624
- 2024
- 3,466
- 2025
- 4,822
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("automated machine learning" OR "AutoML" OR "neural architecture search" OR "Bayesian optimization" OR "active learning" OR "optimal experimental design") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.8k publishing authors. Paper count × 4.24 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,100 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,822 papers (20.7%). 9.4% of authorship records lack an Author ID; 26.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.24; 95% bootstrap interval 4.03–4.49. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.59B.
- AI infrastructure/models/research/governance: $143.2B × 0.35%
- Pharmaceutical: $10.6B × 20.00%
- Biotech: $4.84B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Allocates the search or observation budget rather than optimizing model weights.
Organizations
Selected examples- Preferred Networks Tools & infrastructure
Optuna automates hyperparameter search and experiment selection.
- Meta Tools & infrastructure
Ax and BoTorch provide adaptive experimentation and Bayesian optimization.
- H2O.ai Product
Automated ML and predictive modeling tools.
More organizations (3)
- Citrine Informatics Product
Materials informatics.
- DataRobot Product
Predictive AI and model operations.
- Secondmind Product
Data-efficient engineering optimization.
Data-centric AI
Data curation, labeling, synthetic data, provenance, deduplication, and leakage prevention.
- Authors · scenario
- 7.6k–38k
- Investment · scenario
- $9.58B–$38.3B
Sources and methodology
Publication activity
- 2019
- 734
- 2020
- 1,204
- 2021
- 1,685
- 2022
- 2,252
- 2023
- 3,362
- 2024
- 6,120
- 2025
- 10,400
Data and quality are broad terms, so this query may include substantial noise from applied work.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("data curation" OR "synthetic data" OR "data quality" OR "data cleaning" OR "data annotation" OR "training data selection" OR "data contamination") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 13k publishing authors. Paper count × 3.66 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,642 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 10,400 papers (9.6%). 11.2% of authorship records lack an Author ID; 14.4% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.66; 95% bootstrap interval 3.46–3.89. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $19.2B.
- Data management, processing: $31.6B × 55.00%
- Pharmaceutical: $10.6B × 10.00%
- Biotech: $4.84B × 15.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers datasets across modalities, including language-model training data.
Organizations
Selected examples- Scale AI Data & labeling
Training data and model evaluation.
- Snorkel AI Data & labeling
Programmatic and expert training data.
- Hugging Face Tools & infrastructure
Model and dataset ecosystem.
More organizations (8)
- Appen Data & labeling
Human data and evaluation.
- Labelbox Data & labeling
Data labeling and model training workflows.
- Mercor Data & labeling
Expert training data.
- MOSTLY AI, powered by Syntho Product
Synthetic data generation and privacy-oriented data sharing.
Brand assets owned by Syntho - Prolific Data & labeling
Human data and research participants.
- Roboflow Product
Computer vision development.
- Toloka Data & labeling
Human data and model evaluation.
- Turing Data & labeling
Expert data for model training.