Structured data Sections

Tabular, time-series and graph ML

Learning from tables, time series and graphs, including forecasting and the detection of unusual observations.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Tabular ML

Trees, boosting, neural and foundation models for tables, missing data, imbalance, and structured prediction.

Tabular foundation models Gradient boosting
19,047 papers
3.5× vs. 2023
Authors · scenario
17k–83k
Investment · scenario
$2.88B–$11.5B
Sources and methodology

Publication activity

2019
472
2020
1,177
2021
2,268
2022
3,704
2023
5,481
2024
9,314
2025
19,047

The query includes mentions of boosting, so some applied papers may concern other data types.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("tabular data" OR "tabular learning" OR "tabular foundation model" OR "gradient boosting decision tree" OR "CatBoost" OR "XGBoost" OR "TabPFN") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 28k publishing authors. Paper count × 4.36 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 8,614 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 2,000 of 19,047 papers (10.5%). 9.0% of authorship records lack an Author ID; 23.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.36; 95% bootstrap interval 4.23–4.51. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $5.77B.

  • Data management, processing: $31.6B × 5.00%
  • Retail: $4.10B × 10.00%
  • Medical and healthcare: $11.8B × 15.00%
  • Fintech: $6.52B × 25.00%
  • Accounting/finance: $2.59B × 15.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Unlike time series, rows need not form a temporal sequence.

Organizations

Selected examples
  • Yandex Research Tools & infrastructure

    CatBoost applies gradient-boosted decision trees to structured data.

  • Prior Labs Models

    Tabular foundation models.

  • H2O.ai Product

    Automated ML and predictive modeling tools.

More organizations (2)
  • DataRobot Product

    Predictive AI and model operations.

  • Kumo AI Product

    Graph and relational predictive models.

Time series and forecasting

Forecasting, temporal representations, spatiotemporal models, multivariate series, and change points.

Time-series foundation models Forecasting
2,290 papers
2.0× vs. 2023
Authors · scenario
3.1k–7.4k
Investment · scenario
$3.28B–$13.1B
Sources and methodology

Publication activity

2019
354
2020
588
2021
706
2022
881
2023
1,129
2024
1,513
2025
2,290

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("time series forecasting" OR "time series prediction" OR "time series foundation model" OR "time series representation" OR "spatiotemporal forecasting") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.1k publishing authors. Paper count × 3.25 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,128 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,290 papers (43.7%). 13.2% of authorship records lack an Author ID; 25.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.25; 95% bootstrap interval 3.11–3.39. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $6.57B.

  • Data management, processing: $31.6B × 3.00%
  • Internet of things: $14.6B × 20.00%
  • Fintech: $6.52B × 20.00%
  • Energy management: $4.64B × 30.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Anomaly detection is a related task across time series and other data types.

Organizations

Selected examples
  • Google Research Models

    TimesFM research develops foundation models for time-series forecasting.

  • Nixtla Product

    Time-series foundation models.

  • Amazon / AWS Research

    Published research on probabilistic forecasting and time series.

More organizations (2)
  • Jua Product

    Weather models.

  • PyMC Labs Tools & infrastructure

    PyMC tools support Bayesian forecasting and state-space time-series models.

Graph machine learning

Graph neural networks, graph transformers, link prediction, and geometric or relational representations.

GNNs Graph transformers
11,725 papers
1.9× vs. 2023
Authors · scenario
9.0k–45k
Investment · scenario
$3.46B–$13.9B
Sources and methodology

Publication activity

2019
850
2020
1,848
2021
3,050
2022
4,621
2023
6,271
2024
8,483
2025
11,725

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("graph neural network" OR "graph transformer" OR "graph representation learning" OR "graph machine learning" OR "graph embedding")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 15k publishing authors. Paper count × 3.83 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,723 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 11,725 papers (8.5%). 12.3% of authorship records lack an Author ID; 23.0% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.83; 95% bootstrap interval 3.68–3.98. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $6.93B.

  • Data management, processing: $31.6B × 5.00%
  • Cybersecurity, data protection: $8.42B × 10.00%
  • Pharmaceutical: $10.6B × 25.00%
  • Biotech: $4.84B × 25.00%
  • Fintech: $6.52B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Learning on graphs is distinct from building knowledge graphs.

Organizations

Selected examples
  • Google Research Tools & infrastructure

    TensorFlow GNN supports heterogeneous graph neural networks.

  • Neo4j Product

    Graph data science and knowledge graphs.

  • Kumo AI Product

    Graph and relational predictive models.

More organizations (2)
  • Google DeepMind Research

    GNoME applies graph neural networks to crystal stability prediction.

  • TigerGraph Tools & infrastructure

    Graph machine-learning tools combine entity features with network relationships.

Anomaly and novelty detection

Rare deviations, new classes, and out-of-distribution detection in time series, graphs, tables, and images.

Anomaly detection OOD detection
7,526 papers
2.8× vs. 2023
Authors · scenario
4.5k–23k
Investment · scenario
$3.40B–$13.6B
Sources and methodology

Publication activity

2019
702
2020
1,134
2021
1,670
2022
2,045
2023
2,723
2024
4,185
2025
7,526

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("anomaly detection" OR "novelty detection" OR "out of distribution detection" OR "outlier detection") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 7.6k publishing authors. Paper count × 3.02 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,969 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 7,526 papers (13.3%). 11.5% of authorship records lack an Author ID; 16.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.02; 95% bootstrap interval 2.82–3.28. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $6.81B.

  • Data management, processing: $31.6B × 2.00%
  • Cybersecurity, data protection: $8.42B × 35.00%
  • Internet of things: $14.6B × 10.00%
  • Fintech: $6.52B × 20.00%
  • Energy management: $4.64B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

A cross-cutting task; a paper may also belong to its underlying data modality.

Organizations

Selected examples
  • Nixtla Product

    Time-series foundation models.

  • Feedzai Product

    Financial fraud detection.

  • Darktrace Product

    AI cybersecurity.

More organizations (6)