AI agents and decisions Sections

AI agents and decision-making

Systems that act, use tools and make sequential decisions, from coding agents to reinforcement learning and planning.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Reinforcement learning and bandits

Online and offline RL, exploration, contextual bandits, model-based RL, and policy learning.

RL Offline RL Bandits
29,769 papers
2.0× vs. 2023
Authors · scenario
23k–113k
Investment · scenario
$1.02B–$4.08B
Sources and methodology

Publication activity

2019
5,460
2020
7,958
2021
10,033
2022
11,470
2023
14,744
2024
19,007
2025
29,769

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("reinforcement learning" OR "contextual bandit" OR "multi armed bandit")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 38k publishing authors. Paper count × 3.81 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,752 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 29,769 papers (3.4%). 13.0% of authorship records lack an Author ID; 17.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.81; 95% bootstrap interval 3.58–4.04. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.04B.

  • Robotics: $7.84B × 10.00%
  • Fintech: $6.52B × 5.00%
  • Energy management: $4.64B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

RL used for language-model post-training overlaps with this area.

Organizations

Selected examples
  • Google DeepMind Research

    Models, world models and robotics.

  • OpenAI Research

    Research on learning reward models from human preferences and optimizing policies.

  • Anthropic Research

    Constitutional AI research uses AI feedback in reinforcement learning.

More organizations (4)
  • DeepSeek Research

    DeepSeek-R1 research studies reinforcement learning for reasoning.

  • Google Research Research

    Learning theory, optimization, reinforcement learning and differential privacy research.

  • Sony AI Research

    Gran Turismo Sophy uses deep reinforcement learning for competitive racing in simulation.

  • Spotify Research

    Research applies reinforcement learning to sequential music recommendation.

Multi-agent systems

Learning, communication, negotiation, games, cooperation, and distributed decisions among agents.

Multi-agent learning MARL Coordination
2,322 papers
2.3× vs. 2023
Authors · scenario
3.3k–8.2k
Investment · scenario
$401M–$1.60B
Sources and methodology

Publication activity

2019
292
2020
421
2021
597
2022
765
2023
1,029
2024
1,322
2025
2,322

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("multi agent learning" OR "multi agent reinforcement learning" OR "multi agent coordination" OR "multi agent planning" OR "agent communication")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.3k publishing authors. Paper count × 3.52 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,254 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,322 papers (43.1%). 13.6% of authorship records lack an Author ID; 20.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.52; 95% bootstrap interval 3.37–3.65. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $802M.

  • AI agents: $8.02B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Includes multi-agent RL and agents without language models.

Organizations

Selected examples
  • Google DeepMind Research

    Melting Pot studies cooperation and generalization in multi-agent reinforcement learning.

  • Meta Research

    CICERO studies negotiation, strategic reasoning and cooperation in Diplomacy.

  • Microsoft Tools & infrastructure

    Microsoft Agent Framework provides sequential, concurrent, handoff and group-chat agent orchestration.

More organizations (2)
  • CrewAI Product

    Multi-agent orchestration.

  • OpenAI Tools & infrastructure

    The Agents API supports subagents for delegated tasks.

AI agents and tool use

Tool and computer use, agent memory, multi-step workflows, and autonomous task execution.

Agentic AI Computer use Agent memory
8,339 papers
8.8× vs. 2023
Authors · scenario
7.4k–37k
Investment · scenario
$2.46B–$9.86B
Sources and methodology

Publication activity

2019
20
2020
46
2021
58
2022
144
2023
944
2024
3,040
2025
8,339

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("agent" OR "tool use" OR "tool calling" OR "computer use" OR "autonomous agent") AND (("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 12k publishing authors. Paper count × 4.46 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,294 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 8,339 papers (12.0%). 16.0% of authorship records lack an Author ID; 9.6% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.46; 95% bootstrap interval 4.18–4.77. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $4.93B.

  • AI agents: $8.02B × 55.00%
  • Accounting/finance: $2.59B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on the agent system; base-model properties are covered elsewhere.

Organizations

Selected examples
  • Anthropic Tools & infrastructure

    Claude Agent SDK provides tools, context management and agent execution.

  • OpenAI Tools & infrastructure

    The Agents API provides hosted execution and tools for autonomous workflows.

  • LangChain Product

    Agent frameworks and evaluation.

More organizations (15)
  • AI21 Labs Models

    Language models and enterprise agents.

  • Amazon / AWS Tools & infrastructure

    Foundation-model and agent infrastructure.

  • Cognition Product

    Coding agents.

  • CrewAI Product

    Multi-agent orchestration.

  • Decagon Product

    Customer-service agents.

  • Factory Product

    Software engineering agents.

  • Glean Product

    Enterprise search and agents.

  • Google DeepMind Models

    Gemini supports tool use and multi-step agentic tasks.

  • Harvey Product

    Legal AI.

  • LlamaIndex Product

    Document parsing and agents over enterprise information.

  • Microsoft Tools & infrastructure

    Microsoft Agent Framework provides sequential, concurrent, handoff and group-chat agent orchestration.

  • Mistral AI Models

    Open and enterprise models.

  • Moonshot AI / Kimi Models

    Kimi models and agents for coding and knowledge work.

  • Sierra Product

    Customer-service agents.

  • Writer Product

    Enterprise AI agents.

Coding agents and code generation

Code generation, understanding and repair, software engineering agents, testing, and program synthesis.

Coding agents AI coding Program synthesis
1,805 papers
3.4× vs. 2023
Authors · scenario
3.6k–7.3k
Investment · scenario
$6.04B–$24.1B
Sources and methodology

Publication activity

2019
50
2020
97
2021
125
2022
187
2023
535
2024
1,122
2025
1,805

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("code generation" OR "program synthesis" OR "program repair" OR "software engineering agent" OR "code understanding" OR "code completion") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.6k publishing authors. Paper count × 4.04 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,608 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,805 papers (55.4%). 15.4% of authorship records lack an Author ID; 7.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.04; 95% bootstrap interval 3.81–4.25. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $12.1B.

  • AI infrastructure/models/research/governance: $143.2B × 7.00%
  • AI agents: $8.02B × 15.00%
  • Cybersecurity, data protection: $8.42B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

A task area with software-specific benchmarks and verifiable outcomes.

Organizations

Selected examples
  • Anthropic Tools & infrastructure

    Claude Code and its agent SDK support software engineering tasks.

  • OpenAI Tools & infrastructure

    Codex provides coding agents for software engineering.

  • Anysphere / Cursor Product

    AI coding environment.

More organizations (12)
  • Augment Code Product

    Coding agents.

  • Cognition Product

    Coding agents.

  • Factory Product

    Software engineering agents.

  • Google DeepMind Models

    Gemini models support code generation, debugging and software tasks.

  • Magic Models

    Code models and long context.

  • Meta Models

    Llama models generate code as well as natural-language text.

  • Microsoft Product

    GitHub Copilot provides coding assistance and agents for repository changes and testing.

  • Moonshot AI / Kimi Models

    Kimi models and agents for coding and knowledge work.

  • Replit Product

    Coding agents and app development.

  • Sourcegraph Product

    Code intelligence.

  • Tabnine Product

    Enterprise coding assistance.

    Part of Tricentis
  • Z.ai / GLM Models

    GLM language models and coding assistants.

Planning and symbolic reasoning

State-space search, logical inference, SAT/SMT, constraints, and neuro-symbolic problem solving.

Neuro-symbolic AI Planning SAT / SMT
1,359 papers
2.0× vs. 2023
Authors · scenario
2.6k–3.9k
Investment · scenario
$487M–$1.95B
Sources and methodology

Publication activity

2019
490
2020
648
2021
599
2022
620
2023
664
2024
769
2025
1,359

May include applied and non-ML work on planning and constraints.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("automated planning" OR "classical planning" OR "neurosymbolic reasoning" OR "satisfiability modulo theories" OR "constraint satisfaction" OR "automated reasoning")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 2.6k publishing authors. Paper count × 2.87 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,603 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,359 papers (73.6%). 15.6% of authorship records lack an Author ID; 11.8% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 2.87; 95% bootstrap interval 2.71–3.02. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $974M.

  • AI infrastructure/models/research/governance: $143.2B × 0.40%
  • AI agents: $8.02B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Broader than reasoning within language models. Formal proofs also appear under mathematics.

Organizations

Selected examples
  • Microsoft Tools & infrastructure

    Z3 solves symbolic constraints using satisfiability modulo theories.

  • Imandra Tools & infrastructure

    Automated logical reasoning combines formal specifications with symbolic verification.

  • RelationalAI Product

    Decision agents with semantic models and automated reasoning.

More organizations (2)
  • Shield AI Product

    Autonomous aircraft.

  • Timefold Product

    A constraint-satisfaction solver for scheduling, routing and resource allocation.

Combinatorial optimization

Routing, scheduling, resource allocation, learning to optimize, and heuristic search.

Learning to optimize Scheduling
9,847 papers
3.4× vs. 2023
Authors · scenario
6.7k–33k
Investment · scenario
$1.26B–$5.03B
Sources and methodology

Publication activity

2019
633
2020
1,096
2021
1,562
2022
2,069
2023
2,896
2024
5,129
2025
9,847

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("combinatorial optimization" OR "vehicle routing" OR "job shop scheduling" OR "resource allocation" OR "learning to optimize") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "reinforcement learning")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 11k publishing authors. Paper count × 3.40 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,345 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 9,847 papers (10.2%). 11.1% of authorship records lack an Author ID; 19.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.40; 95% bootstrap interval 3.22–3.58. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.52B.

  • AI infrastructure/models/research/governance: $143.2B × 0.40%
  • Robotics: $7.84B × 10.00%
  • Energy management: $4.64B × 25.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Optimizes an external decision problem rather than neural-network parameters.

Organizations

Selected examples
  • Gurobi Product

    Mathematical optimization.

  • Timefold Product

    A constraint-satisfaction solver for scheduling, routing and resource allocation.

  • Optibus Product

    Transport planning optimization.

More organizations (2)
  • Hexaly Tools & infrastructure

    Optimization solver for routing, scheduling and other discrete decision problems.

  • NextBillion.ai Product

    Route optimization.