AI safety and human interaction Sections

AI safety and human interaction

Evaluating AI systems, understanding their behavior, protecting them and studying how people work with them.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

AI evaluation and hallucinations

Benchmarks, reproducibility, factuality, hallucinations, quality assessment, and measurement validity.

Evals Benchmarks Factuality
37,539 papers
4.6× vs. 2023
Authors · scenario
34k–171k
Investment · scenario
$1.58B–$6.32B
Sources and methodology

Publication activity

2019
404
2020
799
2021
1,367
2022
2,120
2023
8,083
2024
20,224
2025
37,539

The query focuses on language and foundation-model evaluation; it does not cover all evaluation research.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("benchmark" OR "evaluation" OR "hallucination" OR "factuality" OR "reproducibility") AND (("language model" OR "large language model" OR "vision language model") OR "foundation model")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 57k publishing authors. Paper count × 4.55 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 8,847 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 2,000 of 37,539 papers (5.3%). 13.3% of authorship records lack an Author ID; 10.4% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.55; 95% bootstrap interval 4.29–4.81. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.16B.

  • AI infrastructure/models/research/governance: $143.2B × 0.50%
  • Cybersecurity, data protection: $8.42B × 5.00%
  • Medical and healthcare: $11.8B × 10.00%
  • Fintech: $6.52B × 5.00%
  • Legal tech: $3.79B × 10.00%
  • Ed tech: $1.44B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

The category is broad; the publication query mainly captures language and foundation-model evaluation.

Organizations

Selected examples
  • Anthropic Research

    Model safety and interpretability research.

  • Arize AI Product

    Model observability and evaluation.

  • Braintrust Product

    Evaluation and observability.

More organizations (15)
  • Apollo Research Research

    Develops pre-deployment evaluations of strategic deception and other alignment risks.

  • Appen Data & labeling

    Human data and evaluation.

  • Arthur AI Product

    AI evaluation and monitoring.

  • Comet Product

    Experiment tracking and AI evaluation.

  • Databricks Tools & infrastructure

    MLflow supports evaluation and tracing of model and agent behavior.

  • Fiddler AI Product

    Model monitoring, governance and evaluation.

  • Galileo Product

    Agent evaluation and observability.

  • Giskard Product

    AI testing and risk evaluation.

  • Google DeepMind Research

    Research develops holistic evaluations of model behavior and capabilities.

  • Langfuse Product

    Open-source LLM observability.

    Part of ClickHouse
  • Mindgard Product

    AI red teaming.

  • OpenAI Tools & infrastructure

    Evals provides a framework and benchmark registry for language-model evaluation.

  • Scale AI Data & labeling

    Training data and model evaluation.

  • Snorkel AI Data & labeling

    Programmatic and expert training data.

  • Toloka Data & labeling

    Human data and model evaluation.

AI alignment and safety

Scalable oversight, reward hacking, deception, control, and evaluation of dangerous capabilities.

AI safety Scalable oversight Reward hacking
2,172 papers
9.7× vs. 2023
Authors · scenario
1.8k–4.4k
Investment · scenario
$215M–$859M
Sources and methodology

Publication activity

2019
57
2020
50
2021
79
2022
109
2023
224
2024
447
2025
2,172

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("AI alignment" OR "AI safety" OR "scalable oversight" OR "reward hacking" OR "deceptive alignment" OR "AI control")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 1.8k publishing authors. Paper count × 2.02 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 1,798 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,172 papers (46.0%). 27.3% of authorship records lack an Author ID; 8.6% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 2.02; 95% bootstrap interval 1.76–2.33. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $430M.

  • AI infrastructure/models/research/governance: $143.2B × 0.30%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on behavior and controllability. Adversarial attacks and privacy are covered separately.

Organizations

Selected examples
  • Anthropic Research

    Model safety and interpretability research.

  • Apollo Research Research

    Studies model scheming and mitigations for covert pursuit of misaligned goals.

  • OpenAI Research

    Research investigates emergent misalignment and methods for realigning models.

More organizations (1)
  • Google DeepMind Research

    Research evaluates dangerous capabilities and other model safety risks.

Interpretability and explainable AI

Mechanistic interpretability, circuits, sparse autoencoders, attribution, SHAP/LIME, and decision explanations.

Mechanistic interpretability Explainable AI Sparse autoencoders
19,876 papers
4.7× vs. 2023
Authors · scenario
17k–83k
Investment · scenario
$215M–$859M
Sources and methodology

Publication activity

2019
255
2020
660
2021
1,402
2022
2,479
2023
4,206
2024
7,572
2025
19,876

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("mechanistic interpretability" OR "explainable AI" OR "explainable artificial intelligence" OR "feature attribution" OR "SHAP" OR "LIME") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 28k publishing authors. Paper count × 4.16 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,103 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 19,876 papers (5.0%). 10.6% of authorship records lack an Author ID; 22.9% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.16; 95% bootstrap interval 3.99–4.34. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $430M.

  • AI infrastructure/models/research/governance: $143.2B × 0.30%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Includes both model-internal analysis and explanations of model outputs.

Organizations

Selected examples
  • Anthropic Research

    Model safety and interpretability research.

  • Goodfire Tools & infrastructure

    Silico exposes internal model representations to explain and debug behavior.

  • IBM Tools & infrastructure

    IBM-origin AI Explainability 360 provides methods for explaining models and data.

More organizations (2)
  • Google DeepMind Research

    Gemma Scope provides tools for studying internal language-model behavior.

  • OpenAI Research

    Research connects interpretable internal features with model behavior.

Robustness and model security

Adversarial attacks, jailbreaks, prompt injection, poisoning, red teaming, and distribution shifts.

AI security Prompt injection Adversarial ML
3,968 papers
2.0× vs. 2023
Authors · scenario
3.2k–13k
Investment · scenario
$985M–$3.94B
Sources and methodology

Publication activity

2019
763
2020
1,258
2021
1,445
2022
1,584
2023
2,008
2024
2,866
2025
3,968

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("adversarial example" OR "adversarial attack" OR "adversarial robustness" OR "jailbreak" OR "prompt injection" OR "data poisoning" OR "model backdoor") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 4.4k publishing authors. Paper count × 3.33 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,151 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 3,968 papers (25.2%). 16.0% of authorship records lack an Author ID; 13.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.33; 95% bootstrap interval 3.17–3.50. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.97B.

  • AI infrastructure/models/research/governance: $143.2B × 0.20%
  • Cybersecurity, data protection: $8.42B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Protecting AI systems is distinct from using AI for cybersecurity.

Organizations

Selected examples
More organizations (5)

Fairness and responsible AI

Bias, fairness, accountability, auditing, and the social effects of model deployment.

Responsible AI Bias AI auditing
8,586 papers
7.9× vs. 2023
Authors · scenario
4.0k–20k
Investment · scenario
$36M–$143M
Sources and methodology

Publication activity

2019
187
2020
301
2021
436
2022
573
2023
1,083
2024
2,728
2025
8,586

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("algorithmic fairness" OR "algorithmic bias" OR "responsible AI" OR "AI accountability" OR "fair machine learning")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.6k publishing authors. Paper count × 2.32 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,297 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 8,586 papers (11.6%). 16.4% of authorship records lack an Author ID; 12.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 2.32; 95% bootstrap interval 2.17–2.47. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $72M.

  • AI infrastructure/models/research/governance: $143.2B × 0.05%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers research questions around model use; regulation alone is not a learning method.

Organizations

Selected examples
  • Sony AI Research

    FHIBE evaluates demographic and environmental disparities in computer-vision models.

  • IBM Tools & infrastructure

    IBM-origin AI Fairness 360 provides bias measurement and mitigation methods.

  • Holistic AI Product

    Bias assessments test AI systems and support mitigation of identified risks.

More organizations (1)
  • Credo AI Tools & infrastructure

    Credo AI Lens provides a responsible-AI assessment framework, including fairness evaluations.

Privacy, federated learning and unlearning

Differential privacy, learning without centralized data, machine unlearning, and training-data protection.

Differential privacy Federated learning Machine unlearning
6,437 papers
2.2× vs. 2023
Authors · scenario
4.1k–21k
Investment · scenario
$1.59B–$6.37B
Sources and methodology

Publication activity

2019
278
2020
799
2021
1,404
2022
2,011
2023
2,906
2024
4,114
2025
6,437

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("federated learning" OR "machine unlearning" OR "differential privacy" OR "privacy preserving machine learning") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.8k publishing authors. Paper count × 3.19 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,099 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 6,437 papers (15.5%). 12.7% of authorship records lack an Author ID; 15.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.19; 95% bootstrap interval 3.06–3.35. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.19B.

  • Cybersecurity, data protection: $8.42B × 20.00%
  • Medical and healthcare: $11.8B × 10.00%
  • Fintech: $6.52B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Federated learning also has a systems component; these topics are grouped for navigation.

Organizations

Selected examples
  • Apple Machine Learning Research Research

    Research studies differential-privacy guarantees and their utility tradeoffs.

  • Flower Labs Product

    Federated AI.

  • Duality Tools & infrastructure

    Privacy-preserving model collaboration with federated learning and cryptographic tools.

More organizations (2)

Human–AI interaction

Human-in-the-loop systems, collaborative decisions, usability, trust, error correction, and user studies.

Human-in-the-loop Human–AI collaboration
6,342 papers
5.3× vs. 2023
Authors · scenario
3.6k–18k
Investment · scenario
$1.15B–$4.59B
Sources and methodology

Publication activity

2019
316
2020
449
2021
585
2022
830
2023
1,189
2024
1,994
2025
6,342

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("human AI interaction" OR "human AI collaboration" OR "human AI teaming" OR "human in the loop" OR "human centered AI")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.1k publishing authors. Paper count × 2.86 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,795 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 6,342 papers (15.8%). 16.0% of authorship records lack an Author ID; 13.9% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 2.86; 95% bootstrap interval 2.61–3.14. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.29B.

  • AI infrastructure/models/research/governance: $143.2B × 0.30%
  • AI agents: $8.02B × 5.00%
  • Medical and healthcare: $11.8B × 10.00%
  • Ed tech: $1.44B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Studies people working with AI systems, beyond chat interfaces or preference training.

Organizations

Selected examples
  • Microsoft Research

    HAX translates human-AI interaction research into design guidance and evaluation tools.

  • Google Research Research

    People + AI Research studies human-centered AI and develops interaction design tools.

  • Adobe / Firefly Research

    Research studies creative environments that give people control over AI-assisted iteration.

More organizations (1)
  • Meta Research

    CICERO research studies language-based negotiation and cooperation with human players.