ML foundations Sections

Machine learning foundations

The methods behind learning: generalization, representations, optimization, causality, uncertainty and data.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Learning theory

Generalization, sample complexity, statistical learning, and the theory of deep and in-context learning.

Generalization In-context learning
575 papers
1.4× vs. 2023
Authors · observed
1,378
Investment · scenario
$107M–$430M
Sources and methodology

Publication activity

2019
185
2020
281
2021
324
2022
340
2023
397
2024
448
2025
575

The query is restricted to AI and ML; it does not cover all statistical learning theory or mathematical theory.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("statistical learning theory" OR "generalization bound" OR "sample complexity" OR "neural tangent kernel" OR "grokking" OR "theory of deep learning" OR "PAC learning") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

1,378 distinct OpenAlex Author IDs across all 575 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 575 of 575 papers (100.0%). 14.7% of authorship records lack an Author ID; 13.6% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $215M.

  • AI infrastructure/models/research/governance: $143.2B × 0.15%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers learning guarantees. Training algorithms are listed under optimization.

Organizations

Selected examples
  • Google Research Research

    Learning theory, optimization, reinforcement learning and differential privacy research.

  • Microsoft Research

    Microsoft Research studies theoretical foundations of machine learning.

  • Amazon / AWS Research

    Research establishes generalization guarantees for learned ensemble strategies.

Training optimization

Stochastic and distributed optimizers, learning-rate schedules, training stability, and matrix methods.

Optimizers Training stability
2,045 papers
1.8× vs. 2023
Authors · scenario
3.2k–6.7k
Investment · scenario
$430M–$1.72B
Sources and methodology

Publication activity

2019
581
2020
799
2021
995
2022
1,036
2023
1,166
2024
1,460
2025
2,045

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("stochastic gradient descent" OR "neural network optimization" OR "adaptive optimizer" OR "learning rate schedule" OR "sharpness aware minimization" OR "second order optimization") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.2k publishing authors. Paper count × 3.29 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,188 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,045 papers (48.9%). 13.2% of authorship records lack an Author ID; 20.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.29; 95% bootstrap interval 3.13–3.44. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $859M.

  • AI infrastructure/models/research/governance: $143.2B × 0.60%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on learning model parameters. Scheduling and routing problems sit under combinatorial optimization.

Organizations

Selected examples
  • Google Research Research

    Learning theory, optimization, reinforcement learning and differential privacy research.

  • Microsoft Research

    Research on optimization methods for machine learning.

  • Google DeepMind Research

    Optax provides gradient transformations and optimization algorithms.

More organizations (1)
  • Meta Research

    Schedule-free optimization studies training methods that do not require learning-rate decay schedules.

Representation learning

Self-supervised, contrastive and semi-supervised learning, clustering, and dimensionality reduction.

Self-supervised learning Contrastive learning
16,868 papers
1.5× vs. 2023
Authors · scenario
15k–73k
Investment · scenario
$2.11B–$8.46B
Sources and methodology

Publication activity

2019
3,062
2020
4,729
2021
6,676
2022
8,888
2023
11,349
2024
14,113
2025
16,868

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("self supervised learning" OR "representation learning" OR "contrastive learning" OR "semi supervised learning" OR "deep clustering" OR "unsupervised learning")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 24k publishing authors. Paper count × 4.30 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,227 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 16,868 papers (5.9%). 11.8% of authorship records lack an Author ID; 26.0% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.30; 95% bootstrap interval 4.14–4.46. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $4.23B.

  • AI infrastructure/models/research/governance: $143.2B × 1.00%
  • Pharmaceutical: $10.6B × 15.00%
  • Biotech: $4.84B × 25.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Methods used across data types; application areas provide a separate view.

Organizations

Selected examples
  • Meta Research

    DINOv3 provides self-supervised visual representations.

  • OpenAI Research

    CLIP learns visual representations through image-text contrastive training.

  • Google Research Research

    SimCLR research learns visual representations through contrastive training.

More organizations (1)

Transfer and continual learning

Domain adaptation, few-shot and meta-learning, continual learning, and distribution shift.

Continual learning Domain adaptation Few-shot learning
20,365 papers
1.5× vs. 2023
Authors · scenario
16k–81k
Investment · scenario
$358M–$1.43B
Sources and methodology

Publication activity

2019
3,576
2020
6,137
2021
8,757
2022
10,794
2023
13,205
2024
16,094
2025
20,365

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("transfer learning" OR "domain adaptation" OR "continual learning" OR "meta learning" OR "few shot learning" OR "domain generalization")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 27k publishing authors. Paper count × 3.97 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,928 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 20,365 papers (4.9%). 11.3% of authorship records lack an Author ID; 21.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.97; 95% bootstrap interval 3.76–4.23. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $716M.

  • AI infrastructure/models/research/governance: $143.2B × 0.50%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers transferring and retaining knowledge. Model-specific fine-tuning also appears under post-training.

Organizations

Selected examples
  • OpenAI Research

    CLIP research evaluates zero-shot transfer from image-text pretraining.

  • Google Research Research

    Meta-Dataset studies few-shot learning and generalization to unseen datasets.

  • Physical Intelligence Research

    Generalist robot policies learned across tasks and robot embodiments.

More organizations (1)

Probabilistic ML and uncertainty

Bayesian inference, graphical models, Monte Carlo, probabilistic programming, calibration, and conformal prediction.

Bayesian ML Conformal prediction
4,617 papers
1.8× vs. 2023
Authors · scenario
3.9k–19k
Investment · scenario
$1.56B–$6.24B
Sources and methodology

Publication activity

2019
824
2020
1,303
2021
1,713
2022
2,181
2023
2,549
2024
3,202
2025
4,617

Many matches apply probabilistic methods. This count does not measure the size of the methods research community.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("Bayesian neural network" OR "Gaussian process" OR "probabilistic graphical model" OR "probabilistic programming" OR "uncertainty quantification" OR "conformal prediction") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.3k publishing authors. Paper count × 4.11 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,926 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,617 papers (21.7%). 9.2% of authorship records lack an Author ID; 24.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.11; 95% bootstrap interval 3.89–4.38. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.12B.

  • AI infrastructure/models/research/governance: $143.2B × 0.20%
  • Medical and healthcare: $11.8B × 5.00%
  • Pharmaceutical: $10.6B × 10.00%
  • Biotech: $4.84B × 15.00%
  • Energy management: $4.64B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Causal inference is listed separately. Uncertainty research extends beyond language-model hallucinations.

Organizations

Selected examples
  • PyMC Labs Tools & infrastructure

    PyMC Labs develops Bayesian modeling tools and supports probabilistic inference with PyMC.

  • Amazon / AWS Research

    Published research on probabilistic forecasting and time series.

  • Microsoft Tools & infrastructure

    Infer.NET provides probabilistic programming and Bayesian inference.

More organizations (1)
  • Google Research Research

    Research develops automated structured variational inference for probabilistic programs.

Causal inference

Causal discovery, causal representations, intervention effects, and uplift modeling.

Causal ML Causal discovery
1,661 papers
2.8× vs. 2023
Authors · scenario
3.5k–6.3k
Investment · scenario
$1.32B–$5.28B
Sources and methodology

Publication activity

2019
186
2020
286
2021
391
2022
505
2023
588
2024
849
2025
1,661

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("causal inference" OR "causal discovery" OR "causal representation" OR "causal machine learning" OR "treatment effect" OR "uplift modeling") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.5k publishing authors. Paper count × 3.79 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,470 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,661 papers (60.2%). 12.3% of authorship records lack an Author ID; 19.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.79; 95% bootstrap interval 3.53–4.05. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.64B.

  • AI infrastructure/models/research/governance: $143.2B × 0.15%
  • Marketing, digital ads: $2.51B × 5.00%
  • Medical and healthcare: $11.8B × 5.00%
  • Pharmaceutical: $10.6B × 10.00%
  • Fintech: $6.52B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on cause and effect rather than prediction alone.

Organizations

Selected examples
  • Microsoft Research

    Causal machine learning.

  • Uber Tools & infrastructure

    Uber-origin CausalML provides uplift modeling and treatment-effect estimation.

  • PyMC Labs Tools & infrastructure

    CausalPy supports Bayesian causal analysis of quasi-experiments.

More organizations (1)
  • IBM Tools & infrastructure

    Causal 360 provides methods for causal inference and intervention analysis.

AutoML and experimental design

Model selection, hyperparameter search, Bayesian optimization, active learning, and choosing the next experiment.

AutoML Bayesian optimization Active learning
4,822 papers
1.8× vs. 2023
Authors · scenario
4.1k–20k
Investment · scenario
$1.79B–$7.17B
Sources and methodology

Publication activity

2019
751
2020
1,256
2021
1,780
2022
2,026
2023
2,624
2024
3,466
2025
4,822

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("automated machine learning" OR "AutoML" OR "neural architecture search" OR "Bayesian optimization" OR "active learning" OR "optimal experimental design") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.8k publishing authors. Paper count × 4.24 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,100 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,822 papers (20.7%). 9.4% of authorship records lack an Author ID; 26.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.24; 95% bootstrap interval 4.03–4.49. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.59B.

  • AI infrastructure/models/research/governance: $143.2B × 0.35%
  • Pharmaceutical: $10.6B × 20.00%
  • Biotech: $4.84B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Allocates the search or observation budget rather than optimizing model weights.

Organizations

Selected examples
  • Preferred Networks Tools & infrastructure

    Optuna automates hyperparameter search and experiment selection.

  • Meta Tools & infrastructure

    Ax and BoTorch provide adaptive experimentation and Bayesian optimization.

  • H2O.ai Product

    Automated ML and predictive modeling tools.

More organizations (3)

Data-centric AI

Data curation, labeling, synthetic data, provenance, deduplication, and leakage prevention.

Synthetic data Data curation
10,400 papers
3.1× vs. 2023
Authors · scenario
7.6k–38k
Investment · scenario
$9.58B–$38.3B
Sources and methodology

Publication activity

2019
734
2020
1,204
2021
1,685
2022
2,252
2023
3,362
2024
6,120
2025
10,400

Data and quality are broad terms, so this query may include substantial noise from applied work.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("data curation" OR "synthetic data" OR "data quality" OR "data cleaning" OR "data annotation" OR "training data selection" OR "data contamination") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 13k publishing authors. Paper count × 3.66 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,642 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 10,400 papers (9.6%). 11.2% of authorship records lack an Author ID; 14.4% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.66; 95% bootstrap interval 3.46–3.89. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $19.2B.

  • Data management, processing: $31.6B × 55.00%
  • Pharmaceutical: $10.6B × 10.00%
  • Biotech: $4.84B × 15.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers datasets across modalities, including language-model training data.

Organizations

Selected examples
  • Scale AI Data & labeling

    Training data and model evaluation.

  • Snorkel AI Data & labeling

    Programmatic and expert training data.

  • Hugging Face Tools & infrastructure

    Model and dataset ecosystem.

More organizations (8)