AI infrastructure and systems
Training and running models efficiently, from distributed software and inference to chips and on-device deployment.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Distributed training systems
Parallelism, communication, fault tolerance, checkpointing, and compute allocation.
- Authors · observed
- 1,751
- Investment · scenario
- $5.09B–$20.4B
Sources and methodology
Publication activity
- 2019
- 104
- 2020
- 164
- 2021
- 218
- 2022
- 179
- 2023
- 260
- 2024
- 337
- 2025
- 457
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("distributed training" OR "model parallelism" OR "pipeline parallelism" OR "data parallelism" OR "training checkpoint" OR "gradient communication") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
1,751 distinct OpenAlex Author IDs across all 457 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 457 of 457 papers (100.0%). 11.9% of authorship records lack an Author ID; 10.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $10.2B.
- AI infrastructure/models/research/governance: $143.2B × 5.00%
- Cloud computing: $10.3B × 25.00%
- Semiconductors: $4.40B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Optimization algorithms and accelerator hardware are covered in related areas.
Organizations
Selected examples- NVIDIA Tools & infrastructure
GPU systems and distributed computing infrastructure for model training.
- Amazon / AWS Tools & infrastructure
SageMaker AI supports distributed model training on GPUs and Trainium.
- Google Cloud Tools & infrastructure
Cloud TPU systems support large-scale model training.
More organizations (10)
- AMD Tools & infrastructure
AI compute hardware.
- Anyscale Tools & infrastructure
Distributed AI compute and Ray.
- CoreWeave Tools & infrastructure
GPU cloud.
- Crusoe Tools & infrastructure
AI compute infrastructure.
- Flower Labs Product
Federated AI.
- Intel Tools & infrastructure
Gaudi software supports training and fine-tuning large models.
- Lambda Tools & infrastructure
GPU cloud.
- Modal Tools & infrastructure
GPU compute infrastructure.
- Nebius Tools & infrastructure
AI cloud.
- Together AI Tools & infrastructure
Training and inference cloud.
AI inference and serving
Batching, KV caches, speculative decoding, kernels, compilers, and latency–throughput trade-offs.
- Authors · scenario
- 4.2k–5.1k
- Investment · scenario
- $5.97B–$23.9B
Sources and methodology
Publication activity
- 2019
- 45
- 2020
- 55
- 2021
- 40
- 2022
- 53
- 2023
- 115
- 2024
- 479
- 2025
- 1,047
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("LLM serving" OR "inference serving" OR "speculative decoding" OR "KV cache" OR "neural network compiler" OR "GPU kernel" OR "inference optimization")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 4.2k publishing authors. Paper count × 4.86 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,201 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,047 papers (95.5%). 20.8% of authorship records lack an Author ID; 3.5% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.86; 95% bootstrap interval 4.50–5.23. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $11.9B.
- AI infrastructure/models/research/governance: $143.2B × 5.00%
- Cloud computing: $10.3B × 40.00%
- Semiconductors: $4.40B × 15.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on systems implementation; model compression is listed separately.
Organizations
Selected examples- NVIDIA Tools & infrastructure
GPU inference platforms and libraries for model deployment.
- Fireworks AI Tools & infrastructure
Inference and model tuning.
- Baseten Tools & infrastructure
Model inference infrastructure.
More organizations (14)
- Amazon / AWS Tools & infrastructure
Foundation-model and agent infrastructure.
- Anyscale Tools & infrastructure
Distributed AI compute and Ray.
- Cerebras Tools & infrastructure
Wafer-scale AI compute.
- CoreWeave Tools & infrastructure
GPU cloud.
- Google Cloud Tools & infrastructure
Cloud TPU supports model inference workloads.
- Groq Tools & infrastructure
Cloud inference services.
- Hugging Face Tools & infrastructure
Model and dataset ecosystem.
- Inworld AI Product
Speech models and inference APIs.
- Modal Tools & infrastructure
GPU compute infrastructure.
- Nebius Tools & infrastructure
AI cloud.
- Red Hat / Neural Magic Tools & infrastructure
Inference optimization, model quantization and compression.
Part of Red Hat - Replicate Tools & infrastructure
Model inference platform.
Part of Cloudflare - SambaNova Tools & infrastructure
AI inference and model-tuning systems.
- Together AI Tools & infrastructure
Training and inference cloud.
Model compression
Quantization, pruning, distillation, sparsity, and quality–cost trade-offs.
- Authors · scenario
- 3.8k–18k
- Investment · scenario
- $2.37B–$9.47B
Sources and methodology
Publication activity
- 2019
- 432
- 2020
- 802
- 2021
- 1,202
- 2022
- 1,675
- 2023
- 2,344
- 2024
- 3,271
- 2025
- 4,487
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("model quantization" OR "neural network quantization" OR "model pruning" OR "network pruning" OR "knowledge distillation" OR "model compression")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 5.9k publishing authors. Paper count × 3.93 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,813 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,487 papers (22.3%). 13.8% of authorship records lack an Author ID; 21.8% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.93; 95% bootstrap interval 3.76–4.11. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $4.74B.
- AI infrastructure/models/research/governance: $143.2B × 3.00%
- Semiconductors: $4.40B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Applies across model types, including language models.
Organizations
Selected examples- NVIDIA Tools & infrastructure
TensorRT and Model Optimizer support quantization and inference optimization.
- Qualcomm Research
AIMET provides model compression and quantization techniques for efficient deployment.
- DeepSeek Research
DeepSeek-R1 research includes reasoning-model distillation.
More organizations (5)
- Apple Machine Learning Research Research
On-device machine learning research.
- Google DeepMind Models
Gemini research includes distillation into smaller models.
- Latent AI Tools & infrastructure
Edge model optimization.
- Nota AI Tools & infrastructure
Model compression and deployment.
- Red Hat / Neural Magic Tools & infrastructure
Inference optimization, model quantization and compression.
Part of Red Hat
Edge AI and on-device models
Models for mobile and embedded devices with limited memory, energy, and connectivity.
- Authors · scenario
- 3.2k–9.2k
- Investment · scenario
- $3.43B–$13.7B
Sources and methodology
Publication activity
- 2019
- 119
- 2020
- 270
- 2021
- 395
- 2022
- 580
- 2023
- 764
- 2024
- 1,204
- 2025
- 2,732
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("edge AI" OR "edge intelligence" OR "on device learning" OR "tinyML" OR "on device inference" OR "energy efficient inference")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.2k publishing authors. Paper count × 3.38 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,228 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,732 papers (36.6%). 10.8% of authorship records lack an Author ID; 14.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.38; 95% bootstrap interval 3.21–3.55. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $6.86B.
- AI infrastructure/models/research/governance: $143.2B × 1.00%
- Cloud computing: $10.3B × 15.00%
- Internet of things: $14.6B × 25.00%
- Semiconductors: $4.40B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Defined by deployment constraints; compression and privacy are related methods.
Organizations
Selected examples- NVIDIA Tools & infrastructure
TensorRT supports inference on Jetson and DRIVE edge platforms.
- Qualcomm Research
AIMET provides model compression and quantization techniques for efficient deployment.
- Edge Impulse Tools & infrastructure
Embedded ML development.
Part of Qualcomm
More organizations (10)
- Apple Machine Learning Research Research
MobileCLIP develops image-text models with mobile memory and latency constraints.
- Arm Tools & infrastructure
AI compute IP.
- Axelera AI Tools & infrastructure
Edge AI processors.
- Hailo Tools & infrastructure
Edge AI processors.
- Latent AI Tools & infrastructure
Edge model optimization.
- Liquid AI Models
Efficient foundation models.
- Nota AI Tools & infrastructure
Model compression and deployment.
- NXP / Kinara Tools & infrastructure
Neural processors for edge inference.
Part of NXP - Picovoice Tools & infrastructure
Local inference runtimes deploy speech, language and vision models on devices.
- Sensory Product
Compact on-device models recognize sounds without sending recordings to a server.
MLOps and model observability
Data and model pipelines, testing, debugging, monitoring, drift, versioning, and reproducible deployment.
- Authors · scenario
- 3.5k–10k
- Investment · scenario
- $4.19B–$16.8B
Sources and methodology
Publication activity
- 2019
- 300
- 2020
- 537
- 2021
- 727
- 2022
- 971
- 2023
- 1,246
- 2024
- 1,557
- 2025
- 2,827
Papers about applying or deploying models are not always MLOps research.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("MLOps" OR "ML pipeline" OR "machine learning pipeline" OR "model monitoring" OR "data drift" OR "ML testing" OR "machine learning testing" OR "model deployment")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.5k publishing authors. Paper count × 3.67 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,546 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,827 papers (35.4%). 11.0% of authorship records lack an Author ID; 13.8% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.67; 95% bootstrap interval 3.46–3.90. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $8.38B.
- Data management, processing: $31.6B × 20.00%
- Cloud computing: $10.3B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers research on operating ML systems. Routine deployment work is not necessarily research.
Organizations
Selected examples- Databricks Tools & infrastructure
MLflow integrates experiment tracking, model registration and deployment workflows.
- Weights & Biases Product
Experiment tracking and model evaluation.
Part of CoreWeave - Arize AI Product
Model observability and evaluation.
More organizations (8)
- Arthur AI Product
AI evaluation and monitoring.
- Braintrust Product
Evaluation and observability.
- Comet Product
Experiment tracking and AI evaluation.
- DataRobot Product
Predictive AI and model operations.
- Edge Impulse Tools & infrastructure
Embedded ML development.
Part of Qualcomm - Fiddler AI Product
Model monitoring, governance and evaluation.
- Galileo Product
Agent evaluation and observability.
- Langfuse Product
Open-source LLM observability.
Part of ClickHouse
AI hardware and co-design
Accelerators, memory, interconnects, and joint design of algorithms and hardware.
- Authors · observed
- 1,729
- Investment · scenario
- $1.69B–$6.74B
Sources and methodology
Publication activity
- 2019
- 86
- 2020
- 159
- 2021
- 210
- 2022
- 265
- 2023
- 273
- 2024
- 319
- 2025
- 453
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("AI accelerator" OR "neural network accelerator" OR "deep learning accelerator" OR "tensor processing unit" OR "hardware software co design") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
1,729 distinct OpenAlex Author IDs across all 453 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 453 of 453 papers (100.0%). 10.0% of authorship records lack an Author ID; 14.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.37B.
- Internet of things: $14.6B × 5.00%
- Semiconductors: $4.40B × 60.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Hardware is separated from training and inference software where possible.
Organizations
Selected examples- NVIDIA Tools & infrastructure
GPUs and computing systems designed for AI workloads.
- AMD Tools & infrastructure
AI compute hardware.
- Google Cloud Tools & infrastructure
Cloud TPU provides specialized accelerators for machine learning.
More organizations (11)
- Amazon / AWS Tools & infrastructure
AWS Inferentia accelerators are designed for machine learning inference.
- Arm Tools & infrastructure
AI compute IP.
- Axelera AI Tools & infrastructure
Edge AI processors.
- Cerebras Tools & infrastructure
Wafer-scale AI compute.
- Etched Tools & infrastructure
Transformer accelerators.
- FuriosaAI Tools & infrastructure
AI inference processors.
- Hailo Tools & infrastructure
Edge AI processors.
- Intel Tools & infrastructure
Gaudi accelerators target deep learning workloads.
- NXP / Kinara Tools & infrastructure
Neural processors for edge inference.
Part of NXP - Rebellions Tools & infrastructure
AI accelerators.
- Tenstorrent Tools & infrastructure
AI processors.