AI agents and decision-making
Systems that act, use tools and make sequential decisions, from coding agents to reinforcement learning and planning.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Reinforcement learning and bandits
Online and offline RL, exploration, contextual bandits, model-based RL, and policy learning.
- Authors · scenario
- 23k–113k
- Investment · scenario
- $1.02B–$4.08B
Sources and methodology
Publication activity
- 2019
- 5,460
- 2020
- 7,958
- 2021
- 10,033
- 2022
- 11,470
- 2023
- 14,744
- 2024
- 19,007
- 2025
- 29,769
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("reinforcement learning" OR "contextual bandit" OR "multi armed bandit")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 38k publishing authors. Paper count × 3.81 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,752 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 29,769 papers (3.4%). 13.0% of authorship records lack an Author ID; 17.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.81; 95% bootstrap interval 3.58–4.04. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.04B.
- Robotics: $7.84B × 10.00%
- Fintech: $6.52B × 5.00%
- Energy management: $4.64B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
RL used for language-model post-training overlaps with this area.
Organizations
Selected examples- Google DeepMind Research
Models, world models and robotics.
- OpenAI Research
Research on learning reward models from human preferences and optimizing policies.
- Anthropic Research
Constitutional AI research uses AI feedback in reinforcement learning.
More organizations (4)
- DeepSeek Research
DeepSeek-R1 research studies reinforcement learning for reasoning.
- Google Research Research
Learning theory, optimization, reinforcement learning and differential privacy research.
- Sony AI Research
Gran Turismo Sophy uses deep reinforcement learning for competitive racing in simulation.
- Spotify Research
Research applies reinforcement learning to sequential music recommendation.
Multi-agent systems
Learning, communication, negotiation, games, cooperation, and distributed decisions among agents.
- Authors · scenario
- 3.3k–8.2k
- Investment · scenario
- $401M–$1.60B
Sources and methodology
Publication activity
- 2019
- 292
- 2020
- 421
- 2021
- 597
- 2022
- 765
- 2023
- 1,029
- 2024
- 1,322
- 2025
- 2,322
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("multi agent learning" OR "multi agent reinforcement learning" OR "multi agent coordination" OR "multi agent planning" OR "agent communication")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.3k publishing authors. Paper count × 3.52 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,254 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,322 papers (43.1%). 13.6% of authorship records lack an Author ID; 20.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.52; 95% bootstrap interval 3.37–3.65. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $802M.
- AI agents: $8.02B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Includes multi-agent RL and agents without language models.
Organizations
Selected examples- Google DeepMind Research
Melting Pot studies cooperation and generalization in multi-agent reinforcement learning.
- Meta Research
CICERO studies negotiation, strategic reasoning and cooperation in Diplomacy.
- Microsoft Tools & infrastructure
Microsoft Agent Framework provides sequential, concurrent, handoff and group-chat agent orchestration.
AI agents and tool use
Tool and computer use, agent memory, multi-step workflows, and autonomous task execution.
- Authors · scenario
- 7.4k–37k
- Investment · scenario
- $2.46B–$9.86B
Sources and methodology
Publication activity
- 2019
- 20
- 2020
- 46
- 2021
- 58
- 2022
- 144
- 2023
- 944
- 2024
- 3,040
- 2025
- 8,339
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("agent" OR "tool use" OR "tool calling" OR "computer use" OR "autonomous agent") AND (("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 12k publishing authors. Paper count × 4.46 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,294 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 8,339 papers (12.0%). 16.0% of authorship records lack an Author ID; 9.6% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.46; 95% bootstrap interval 4.18–4.77. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $4.93B.
- AI agents: $8.02B × 55.00%
- Accounting/finance: $2.59B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on the agent system; base-model properties are covered elsewhere.
Organizations
Selected examples- Anthropic Tools & infrastructure
Claude Agent SDK provides tools, context management and agent execution.
- OpenAI Tools & infrastructure
The Agents API provides hosted execution and tools for autonomous workflows.
- LangChain Product
Agent frameworks and evaluation.
More organizations (15)
- AI21 Labs Models
Language models and enterprise agents.
- Amazon / AWS Tools & infrastructure
Foundation-model and agent infrastructure.
- Cognition Product
Coding agents.
- CrewAI Product
Multi-agent orchestration.
- Decagon Product
Customer-service agents.
- Factory Product
Software engineering agents.
- Glean Product
Enterprise search and agents.
- Google DeepMind Models
Gemini supports tool use and multi-step agentic tasks.
- Harvey Product
Legal AI.
- LlamaIndex Product
Document parsing and agents over enterprise information.
- Microsoft Tools & infrastructure
Microsoft Agent Framework provides sequential, concurrent, handoff and group-chat agent orchestration.
- Mistral AI Models
Open and enterprise models.
- Moonshot AI / Kimi Models
Kimi models and agents for coding and knowledge work.
- Sierra Product
Customer-service agents.
- Writer Product
Enterprise AI agents.
Coding agents and code generation
Code generation, understanding and repair, software engineering agents, testing, and program synthesis.
- Authors · scenario
- 3.6k–7.3k
- Investment · scenario
- $6.04B–$24.1B
Sources and methodology
Publication activity
- 2019
- 50
- 2020
- 97
- 2021
- 125
- 2022
- 187
- 2023
- 535
- 2024
- 1,122
- 2025
- 1,805
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("code generation" OR "program synthesis" OR "program repair" OR "software engineering agent" OR "code understanding" OR "code completion") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.6k publishing authors. Paper count × 4.04 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,608 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,805 papers (55.4%). 15.4% of authorship records lack an Author ID; 7.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.04; 95% bootstrap interval 3.81–4.25. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $12.1B.
- AI infrastructure/models/research/governance: $143.2B × 7.00%
- AI agents: $8.02B × 15.00%
- Cybersecurity, data protection: $8.42B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
A task area with software-specific benchmarks and verifiable outcomes.
Organizations
Selected examples- Anthropic Tools & infrastructure
Claude Code and its agent SDK support software engineering tasks.
- OpenAI Tools & infrastructure
Codex provides coding agents for software engineering.
- Anysphere / Cursor Product
AI coding environment.
More organizations (12)
- Augment Code Product
Coding agents.
- Cognition Product
Coding agents.
- Factory Product
Software engineering agents.
- Google DeepMind Models
Gemini models support code generation, debugging and software tasks.
- Magic Models
Code models and long context.
- Meta Models
Llama models generate code as well as natural-language text.
- Microsoft Product
GitHub Copilot provides coding assistance and agents for repository changes and testing.
- Moonshot AI / Kimi Models
Kimi models and agents for coding and knowledge work.
- Replit Product
Coding agents and app development.
- Sourcegraph Product
Code intelligence.
- Tabnine Product
Enterprise coding assistance.
Part of Tricentis - Z.ai / GLM Models
GLM language models and coding assistants.
Planning and symbolic reasoning
State-space search, logical inference, SAT/SMT, constraints, and neuro-symbolic problem solving.
- Authors · scenario
- 2.6k–3.9k
- Investment · scenario
- $487M–$1.95B
Sources and methodology
Publication activity
- 2019
- 490
- 2020
- 648
- 2021
- 599
- 2022
- 620
- 2023
- 664
- 2024
- 769
- 2025
- 1,359
May include applied and non-ML work on planning and constraints.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("automated planning" OR "classical planning" OR "neurosymbolic reasoning" OR "satisfiability modulo theories" OR "constraint satisfaction" OR "automated reasoning")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 2.6k publishing authors. Paper count × 2.87 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,603 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,359 papers (73.6%). 15.6% of authorship records lack an Author ID; 11.8% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 2.87; 95% bootstrap interval 2.71–3.02. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $974M.
- AI infrastructure/models/research/governance: $143.2B × 0.40%
- AI agents: $8.02B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Broader than reasoning within language models. Formal proofs also appear under mathematics.
Organizations
Selected examples- Microsoft Tools & infrastructure
Z3 solves symbolic constraints using satisfiability modulo theories.
- Imandra Tools & infrastructure
Automated logical reasoning combines formal specifications with symbolic verification.
- RelationalAI Product
Decision agents with semantic models and automated reasoning.
Combinatorial optimization
Routing, scheduling, resource allocation, learning to optimize, and heuristic search.
- Authors · scenario
- 6.7k–33k
- Investment · scenario
- $1.26B–$5.03B
Sources and methodology
Publication activity
- 2019
- 633
- 2020
- 1,096
- 2021
- 1,562
- 2022
- 2,069
- 2023
- 2,896
- 2024
- 5,129
- 2025
- 9,847
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("combinatorial optimization" OR "vehicle routing" OR "job shop scheduling" OR "resource allocation" OR "learning to optimize") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "reinforcement learning")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 11k publishing authors. Paper count × 3.40 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,345 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 9,847 papers (10.2%). 11.1% of authorship records lack an Author ID; 19.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.40; 95% bootstrap interval 3.22–3.58. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.52B.
- AI infrastructure/models/research/governance: $143.2B × 0.40%
- Robotics: $7.84B × 10.00%
- Energy management: $4.64B × 25.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Optimizes an external decision problem rather than neural-network parameters.
Organizations
Selected examples- Gurobi Product
Mathematical optimization.
- Timefold Product
A constraint-satisfaction solver for scheduling, routing and resource allocation.
- Optibus Product
Transport planning optimization.
More organizations (2)
- Hexaly Tools & infrastructure
Optimization solver for routing, scheduling and other discrete decision problems.
- NextBillion.ai Product
Route optimization.