AI safety and human interaction
Evaluating AI systems, understanding their behavior, protecting them and studying how people work with them.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
AI evaluation and hallucinations
Benchmarks, reproducibility, factuality, hallucinations, quality assessment, and measurement validity.
- Authors · scenario
- 34k–171k
- Investment · scenario
- $1.58B–$6.32B
Sources and methodology
Publication activity
- 2019
- 404
- 2020
- 799
- 2021
- 1,367
- 2022
- 2,120
- 2023
- 8,083
- 2024
- 20,224
- 2025
- 37,539
The query focuses on language and foundation-model evaluation; it does not cover all evaluation research.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("benchmark" OR "evaluation" OR "hallucination" OR "factuality" OR "reproducibility") AND (("language model" OR "large language model" OR "vision language model") OR "foundation model")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 57k publishing authors. Paper count × 4.55 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 8,847 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 2,000 of 37,539 papers (5.3%). 13.3% of authorship records lack an Author ID; 10.4% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.55; 95% bootstrap interval 4.29–4.81. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.16B.
- AI infrastructure/models/research/governance: $143.2B × 0.50%
- Cybersecurity, data protection: $8.42B × 5.00%
- Medical and healthcare: $11.8B × 10.00%
- Fintech: $6.52B × 5.00%
- Legal tech: $3.79B × 10.00%
- Ed tech: $1.44B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
The category is broad; the publication query mainly captures language and foundation-model evaluation.
Organizations
Selected examples- Anthropic Research
Model safety and interpretability research.
- Arize AI Product
Model observability and evaluation.
- Braintrust Product
Evaluation and observability.
More organizations (15)
- Apollo Research Research
Develops pre-deployment evaluations of strategic deception and other alignment risks.
- Appen Data & labeling
Human data and evaluation.
- Arthur AI Product
AI evaluation and monitoring.
- Comet Product
Experiment tracking and AI evaluation.
- Databricks Tools & infrastructure
MLflow supports evaluation and tracing of model and agent behavior.
- Fiddler AI Product
Model monitoring, governance and evaluation.
- Galileo Product
Agent evaluation and observability.
- Giskard Product
AI testing and risk evaluation.
- Google DeepMind Research
Research develops holistic evaluations of model behavior and capabilities.
- Langfuse Product
Open-source LLM observability.
Part of ClickHouse - Mindgard Product
AI red teaming.
- OpenAI Tools & infrastructure
Evals provides a framework and benchmark registry for language-model evaluation.
- Scale AI Data & labeling
Training data and model evaluation.
- Snorkel AI Data & labeling
Programmatic and expert training data.
- Toloka Data & labeling
Human data and model evaluation.
AI alignment and safety
Scalable oversight, reward hacking, deception, control, and evaluation of dangerous capabilities.
- Authors · scenario
- 1.8k–4.4k
- Investment · scenario
- $215M–$859M
Sources and methodology
Publication activity
- 2019
- 57
- 2020
- 50
- 2021
- 79
- 2022
- 109
- 2023
- 224
- 2024
- 447
- 2025
- 2,172
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("AI alignment" OR "AI safety" OR "scalable oversight" OR "reward hacking" OR "deceptive alignment" OR "AI control")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 1.8k publishing authors. Paper count × 2.02 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 1,798 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,172 papers (46.0%). 27.3% of authorship records lack an Author ID; 8.6% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 2.02; 95% bootstrap interval 1.76–2.33. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $430M.
- AI infrastructure/models/research/governance: $143.2B × 0.30%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on behavior and controllability. Adversarial attacks and privacy are covered separately.
Organizations
Selected examples- Anthropic Research
Model safety and interpretability research.
- Apollo Research Research
Studies model scheming and mitigations for covert pursuit of misaligned goals.
- OpenAI Research
Research investigates emergent misalignment and methods for realigning models.
More organizations (1)
- Google DeepMind Research
Research evaluates dangerous capabilities and other model safety risks.
Interpretability and explainable AI
Mechanistic interpretability, circuits, sparse autoencoders, attribution, SHAP/LIME, and decision explanations.
- Authors · scenario
- 17k–83k
- Investment · scenario
- $215M–$859M
Sources and methodology
Publication activity
- 2019
- 255
- 2020
- 660
- 2021
- 1,402
- 2022
- 2,479
- 2023
- 4,206
- 2024
- 7,572
- 2025
- 19,876
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("mechanistic interpretability" OR "explainable AI" OR "explainable artificial intelligence" OR "feature attribution" OR "SHAP" OR "LIME") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 28k publishing authors. Paper count × 4.16 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,103 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 19,876 papers (5.0%). 10.6% of authorship records lack an Author ID; 22.9% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.16; 95% bootstrap interval 3.99–4.34. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $430M.
- AI infrastructure/models/research/governance: $143.2B × 0.30%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Includes both model-internal analysis and explanations of model outputs.
Organizations
Selected examples- Anthropic Research
Model safety and interpretability research.
- Goodfire Tools & infrastructure
Silico exposes internal model representations to explain and debug behavior.
- IBM Tools & infrastructure
IBM-origin AI Explainability 360 provides methods for explaining models and data.
More organizations (2)
- Google DeepMind Research
Gemma Scope provides tools for studying internal language-model behavior.
- OpenAI Research
Research connects interpretable internal features with model behavior.
Robustness and model security
Adversarial attacks, jailbreaks, prompt injection, poisoning, red teaming, and distribution shifts.
- Authors · scenario
- 3.2k–13k
- Investment · scenario
- $985M–$3.94B
Sources and methodology
Publication activity
- 2019
- 763
- 2020
- 1,258
- 2021
- 1,445
- 2022
- 1,584
- 2023
- 2,008
- 2024
- 2,866
- 2025
- 3,968
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("adversarial example" OR "adversarial attack" OR "adversarial robustness" OR "jailbreak" OR "prompt injection" OR "data poisoning" OR "model backdoor") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 4.4k publishing authors. Paper count × 3.33 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,151 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 3,968 papers (25.2%). 16.0% of authorship records lack an Author ID; 13.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.33; 95% bootstrap interval 3.17–3.50. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.97B.
- AI infrastructure/models/research/governance: $143.2B × 0.20%
- Cybersecurity, data protection: $8.42B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Protecting AI systems is distinct from using AI for cybersecurity.
Organizations
Selected examples- HiddenLayer Product
AI security.
- Lakera Product
LLM and agent security.
Part of Check Point - Protect AI Product
AI system security.
Part of Palo Alto Networks
More organizations (5)
- Cisco / Robust Intelligence Product
Protection and testing of AI applications.
Part of Cisco - Giskard Product
AI testing and risk evaluation.
- IBM Tools & infrastructure
IBM-origin Adversarial Robustness Toolbox supports attacks and defenses for ML models.
- Mindgard Product
AI red teaming.
- Resemble AI Product
Deepfake detection and media verification.
Fairness and responsible AI
Bias, fairness, accountability, auditing, and the social effects of model deployment.
- Authors · scenario
- 4.0k–20k
- Investment · scenario
- $36M–$143M
Sources and methodology
Publication activity
- 2019
- 187
- 2020
- 301
- 2021
- 436
- 2022
- 573
- 2023
- 1,083
- 2024
- 2,728
- 2025
- 8,586
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("algorithmic fairness" OR "algorithmic bias" OR "responsible AI" OR "AI accountability" OR "fair machine learning")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.6k publishing authors. Paper count × 2.32 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,297 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 8,586 papers (11.6%). 16.4% of authorship records lack an Author ID; 12.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 2.32; 95% bootstrap interval 2.17–2.47. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $72M.
- AI infrastructure/models/research/governance: $143.2B × 0.05%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers research questions around model use; regulation alone is not a learning method.
Organizations
Selected examples- Sony AI Research
FHIBE evaluates demographic and environmental disparities in computer-vision models.
- IBM Tools & infrastructure
IBM-origin AI Fairness 360 provides bias measurement and mitigation methods.
- Holistic AI Product
Bias assessments test AI systems and support mitigation of identified risks.
More organizations (1)
- Credo AI Tools & infrastructure
Credo AI Lens provides a responsible-AI assessment framework, including fairness evaluations.
Privacy, federated learning and unlearning
Differential privacy, learning without centralized data, machine unlearning, and training-data protection.
- Authors · scenario
- 4.1k–21k
- Investment · scenario
- $1.59B–$6.37B
Sources and methodology
Publication activity
- 2019
- 278
- 2020
- 799
- 2021
- 1,404
- 2022
- 2,011
- 2023
- 2,906
- 2024
- 4,114
- 2025
- 6,437
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("federated learning" OR "machine unlearning" OR "differential privacy" OR "privacy preserving machine learning") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.8k publishing authors. Paper count × 3.19 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,099 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 6,437 papers (15.5%). 12.7% of authorship records lack an Author ID; 15.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.19; 95% bootstrap interval 3.06–3.35. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.19B.
- Cybersecurity, data protection: $8.42B × 20.00%
- Medical and healthcare: $11.8B × 10.00%
- Fintech: $6.52B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Federated learning also has a systems component; these topics are grouped for navigation.
Organizations
Selected examples- Apple Machine Learning Research Research
Research studies differential-privacy guarantees and their utility tradeoffs.
- Flower Labs Product
Federated AI.
- Duality Tools & infrastructure
Privacy-preserving model collaboration with federated learning and cryptographic tools.
More organizations (2)
- Google Research Research
Learning theory, optimization, reinforcement learning and differential privacy research.
- MOSTLY AI, powered by Syntho Product
Synthetic data generation and privacy-oriented data sharing.
Brand assets owned by Syntho
Human–AI interaction
Human-in-the-loop systems, collaborative decisions, usability, trust, error correction, and user studies.
- Authors · scenario
- 3.6k–18k
- Investment · scenario
- $1.15B–$4.59B
Sources and methodology
Publication activity
- 2019
- 316
- 2020
- 449
- 2021
- 585
- 2022
- 830
- 2023
- 1,189
- 2024
- 1,994
- 2025
- 6,342
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("human AI interaction" OR "human AI collaboration" OR "human AI teaming" OR "human in the loop" OR "human centered AI")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.1k publishing authors. Paper count × 2.86 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,795 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 6,342 papers (15.8%). 16.0% of authorship records lack an Author ID; 13.9% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 2.86; 95% bootstrap interval 2.61–3.14. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.29B.
- AI infrastructure/models/research/governance: $143.2B × 0.30%
- AI agents: $8.02B × 5.00%
- Medical and healthcare: $11.8B × 10.00%
- Ed tech: $1.44B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Studies people working with AI systems, beyond chat interfaces or preference training.
Organizations
Selected examples- Microsoft Research
HAX translates human-AI interaction research into design guidance and evaluation tools.
- Google Research Research
People + AI Research studies human-centered AI and develops interaction design tools.
- Adobe / Firefly Research
Research studies creative environments that give people control over AI-assisted iteration.
More organizations (1)
- Meta Research
CICERO research studies language-based negotiation and cooperation with human players.