AI infrastructure Sections

AI infrastructure and systems

Training and running models efficiently, from distributed software and inference to chips and on-device deployment.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Distributed training systems

Parallelism, communication, fault tolerance, checkpointing, and compute allocation.

Distributed training GPU clusters
457 papers
1.8× vs. 2023
Authors · observed
1,751
Investment · scenario
$5.09B–$20.4B
Sources and methodology

Publication activity

2019
104
2020
164
2021
218
2022
179
2023
260
2024
337
2025
457

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("distributed training" OR "model parallelism" OR "pipeline parallelism" OR "data parallelism" OR "training checkpoint" OR "gradient communication") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

1,751 distinct OpenAlex Author IDs across all 457 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 457 of 457 papers (100.0%). 11.9% of authorship records lack an Author ID; 10.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $10.2B.

  • AI infrastructure/models/research/governance: $143.2B × 5.00%
  • Cloud computing: $10.3B × 25.00%
  • Semiconductors: $4.40B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Optimization algorithms and accelerator hardware are covered in related areas.

Organizations

Selected examples
  • NVIDIA Tools & infrastructure

    GPU systems and distributed computing infrastructure for model training.

  • Amazon / AWS Tools & infrastructure

    SageMaker AI supports distributed model training on GPUs and Trainium.

  • Google Cloud Tools & infrastructure

    Cloud TPU systems support large-scale model training.

More organizations (10)
  • AMD Tools & infrastructure

    AI compute hardware.

  • Anyscale Tools & infrastructure

    Distributed AI compute and Ray.

  • CoreWeave Tools & infrastructure

    GPU cloud.

  • Crusoe Tools & infrastructure

    AI compute infrastructure.

  • Flower Labs Product

    Federated AI.

  • Intel Tools & infrastructure

    Gaudi software supports training and fine-tuning large models.

  • Lambda Tools & infrastructure

    GPU cloud.

  • Modal Tools & infrastructure

    GPU compute infrastructure.

  • Nebius Tools & infrastructure

    AI cloud.

  • Together AI Tools & infrastructure

    Training and inference cloud.

AI inference and serving

Batching, KV caches, speculative decoding, kernels, compilers, and latency–throughput trade-offs.

Inference infrastructure Speculative decoding KV cache
1,047 papers
9.1× vs. 2023
Authors · scenario
4.2k–5.1k
Investment · scenario
$5.97B–$23.9B
Sources and methodology

Publication activity

2019
45
2020
55
2021
40
2022
53
2023
115
2024
479
2025
1,047

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("LLM serving" OR "inference serving" OR "speculative decoding" OR "KV cache" OR "neural network compiler" OR "GPU kernel" OR "inference optimization")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 4.2k publishing authors. Paper count × 4.86 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,201 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,047 papers (95.5%). 20.8% of authorship records lack an Author ID; 3.5% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.86; 95% bootstrap interval 4.50–5.23. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $11.9B.

  • AI infrastructure/models/research/governance: $143.2B × 5.00%
  • Cloud computing: $10.3B × 40.00%
  • Semiconductors: $4.40B × 15.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on systems implementation; model compression is listed separately.

Organizations

Selected examples
  • NVIDIA Tools & infrastructure

    GPU inference platforms and libraries for model deployment.

  • Fireworks AI Tools & infrastructure

    Inference and model tuning.

  • Baseten Tools & infrastructure

    Model inference infrastructure.

More organizations (14)
  • Amazon / AWS Tools & infrastructure

    Foundation-model and agent infrastructure.

  • Anyscale Tools & infrastructure

    Distributed AI compute and Ray.

  • Cerebras Tools & infrastructure

    Wafer-scale AI compute.

  • CoreWeave Tools & infrastructure

    GPU cloud.

  • Google Cloud Tools & infrastructure

    Cloud TPU supports model inference workloads.

  • Groq Tools & infrastructure

    Cloud inference services.

  • Hugging Face Tools & infrastructure

    Model and dataset ecosystem.

  • Inworld AI Product

    Speech models and inference APIs.

  • Modal Tools & infrastructure

    GPU compute infrastructure.

  • Nebius Tools & infrastructure

    AI cloud.

  • Red Hat / Neural Magic Tools & infrastructure

    Inference optimization, model quantization and compression.

    Part of Red Hat
  • Replicate Tools & infrastructure

    Model inference platform.

    Part of Cloudflare
  • SambaNova Tools & infrastructure

    AI inference and model-tuning systems.

  • Together AI Tools & infrastructure

    Training and inference cloud.

Model compression

Quantization, pruning, distillation, sparsity, and quality–cost trade-offs.

Quantization Distillation Pruning
4,487 papers
1.9× vs. 2023
Authors · scenario
3.8k–18k
Investment · scenario
$2.37B–$9.47B
Sources and methodology

Publication activity

2019
432
2020
802
2021
1,202
2022
1,675
2023
2,344
2024
3,271
2025
4,487

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("model quantization" OR "neural network quantization" OR "model pruning" OR "network pruning" OR "knowledge distillation" OR "model compression")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 5.9k publishing authors. Paper count × 3.93 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,813 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,487 papers (22.3%). 13.8% of authorship records lack an Author ID; 21.8% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.93; 95% bootstrap interval 3.76–4.11. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $4.74B.

  • AI infrastructure/models/research/governance: $143.2B × 3.00%
  • Semiconductors: $4.40B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Applies across model types, including language models.

Organizations

Selected examples
  • NVIDIA Tools & infrastructure

    TensorRT and Model Optimizer support quantization and inference optimization.

  • Qualcomm Research

    AIMET provides model compression and quantization techniques for efficient deployment.

  • DeepSeek Research

    DeepSeek-R1 research includes reasoning-model distillation.

More organizations (5)

Edge AI and on-device models

Models for mobile and embedded devices with limited memory, energy, and connectivity.

On-device AI TinyML
2,732 papers
3.6× vs. 2023
Authors · scenario
3.2k–9.2k
Investment · scenario
$3.43B–$13.7B
Sources and methodology

Publication activity

2019
119
2020
270
2021
395
2022
580
2023
764
2024
1,204
2025
2,732

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("edge AI" OR "edge intelligence" OR "on device learning" OR "tinyML" OR "on device inference" OR "energy efficient inference")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.2k publishing authors. Paper count × 3.38 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,228 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,732 papers (36.6%). 10.8% of authorship records lack an Author ID; 14.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.38; 95% bootstrap interval 3.21–3.55. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $6.86B.

  • AI infrastructure/models/research/governance: $143.2B × 1.00%
  • Cloud computing: $10.3B × 15.00%
  • Internet of things: $14.6B × 25.00%
  • Semiconductors: $4.40B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Defined by deployment constraints; compression and privacy are related methods.

Organizations

Selected examples
  • NVIDIA Tools & infrastructure

    TensorRT supports inference on Jetson and DRIVE edge platforms.

  • Qualcomm Research

    AIMET provides model compression and quantization techniques for efficient deployment.

  • Edge Impulse Tools & infrastructure

    Embedded ML development.

    Part of Qualcomm
More organizations (10)
  • Apple Machine Learning Research Research

    MobileCLIP develops image-text models with mobile memory and latency constraints.

  • Arm Tools & infrastructure

    AI compute IP.

  • Axelera AI Tools & infrastructure

    Edge AI processors.

  • Hailo Tools & infrastructure

    Edge AI processors.

  • Latent AI Tools & infrastructure

    Edge model optimization.

  • Liquid AI Models

    Efficient foundation models.

  • Nota AI Tools & infrastructure

    Model compression and deployment.

  • NXP / Kinara Tools & infrastructure

    Neural processors for edge inference.

    Part of NXP
  • Picovoice Tools & infrastructure

    Local inference runtimes deploy speech, language and vision models on devices.

  • Sensory Product

    Compact on-device models recognize sounds without sending recordings to a server.

MLOps and model observability

Data and model pipelines, testing, debugging, monitoring, drift, versioning, and reproducible deployment.

Model monitoring Data drift
2,827 papers
2.3× vs. 2023
Authors · scenario
3.5k–10k
Investment · scenario
$4.19B–$16.8B
Sources and methodology

Publication activity

2019
300
2020
537
2021
727
2022
971
2023
1,246
2024
1,557
2025
2,827

Papers about applying or deploying models are not always MLOps research.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("MLOps" OR "ML pipeline" OR "machine learning pipeline" OR "model monitoring" OR "data drift" OR "ML testing" OR "machine learning testing" OR "model deployment")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.5k publishing authors. Paper count × 3.67 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,546 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,827 papers (35.4%). 11.0% of authorship records lack an Author ID; 13.8% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.67; 95% bootstrap interval 3.46–3.90. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $8.38B.

  • Data management, processing: $31.6B × 20.00%
  • Cloud computing: $10.3B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers research on operating ML systems. Routine deployment work is not necessarily research.

Organizations

Selected examples
More organizations (8)

AI hardware and co-design

Accelerators, memory, interconnects, and joint design of algorithms and hardware.

AI accelerators AI chips Hardware–software co-design
453 papers
1.7× vs. 2023
Authors · observed
1,729
Investment · scenario
$1.69B–$6.74B
Sources and methodology

Publication activity

2019
86
2020
159
2021
210
2022
265
2023
273
2024
319
2025
453

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("AI accelerator" OR "neural network accelerator" OR "deep learning accelerator" OR "tensor processing unit" OR "hardware software co design") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

1,729 distinct OpenAlex Author IDs across all 453 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 453 of 453 papers (100.0%). 10.0% of authorship records lack an Author ID; 14.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.37B.

  • Internet of things: $14.6B × 5.00%
  • Semiconductors: $4.40B × 60.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Hardware is separated from training and inference software where possible.

Organizations

Selected examples
  • NVIDIA Tools & infrastructure

    GPUs and computing systems designed for AI workloads.

  • AMD Tools & infrastructure

    AI compute hardware.

  • Google Cloud Tools & infrastructure

    Cloud TPU provides specialized accelerators for machine learning.

More organizations (11)
  • Amazon / AWS Tools & infrastructure

    AWS Inferentia accelerators are designed for machine learning inference.

  • Arm Tools & infrastructure

    AI compute IP.

  • Axelera AI Tools & infrastructure

    Edge AI processors.

  • Cerebras Tools & infrastructure

    Wafer-scale AI compute.

  • Etched Tools & infrastructure

    Transformer accelerators.

  • FuriosaAI Tools & infrastructure

    AI inference processors.

  • Hailo Tools & infrastructure

    Edge AI processors.

  • Intel Tools & infrastructure

    Gaudi accelerators target deep learning workloads.

  • NXP / Kinara Tools & infrastructure

    Neural processors for edge inference.

    Part of NXP
  • Rebellions Tools & infrastructure

    AI accelerators.

  • Tenstorrent Tools & infrastructure

    AI processors.