Natural language processing
Working with language across translation, knowledge extraction, dialogue and text understanding.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Translation and multilingual NLP
Machine translation, cross-lingual transfer, language coverage, and low-resource methods.
- Authors · scenario
- 2.9k–13k
- Investment · scenario
- $681M–$2.72B
Sources and methodology
Publication activity
- 2019
- 1,574
- 2020
- 1,969
- 2021
- 2,203
- 2022
- 2,151
- 2023
- 2,688
- 2024
- 3,376
- 2025
- 4,357
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("machine translation" OR "multilingual language model" OR "cross lingual" OR "low resource language")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 4.5k publishing authors. Paper count × 3.07 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,912 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,357 papers (23.0%). 12.3% of authorship records lack an Author ID; 10.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.07; 95% bootstrap interval 2.89–3.27. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.36B.
- AI infrastructure/models/research/governance: $143.2B × 0.80%
- Ed tech: $1.44B × 15.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Language-specific quality and translation remain distinct tasks alongside general pre-training.
Organizations
Selected examples- DeepL Product
Machine translation.
- Google Cloud Tools & infrastructure
Cloud Translation offers neural and language-model translation.
- Cohere Models
Rerank models rank relevant documents across languages.
Information extraction and knowledge graphs
Entities, relations, events, knowledge graphs, ontologies, and neuro-symbolic knowledge representation.
- Authors · scenario
- 4.2k–21k
- Investment · scenario
- $1.64B–$6.57B
Sources and methodology
Publication activity
- 2019
- 868
- 2020
- 1,278
- 2021
- 1,526
- 2022
- 1,978
- 2023
- 2,793
- 2024
- 3,915
- 2025
- 5,423
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("information extraction" OR "named entity recognition" OR "relation extraction" OR "knowledge graph" OR "knowledge representation" OR "neurosymbolic") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model") OR "natural language processing")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.9k publishing authors. Paper count × 3.84 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,724 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 5,423 papers (18.4%). 13.1% of authorship records lack an Author ID; 17.9% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.84; 95% bootstrap interval 3.60–4.09. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $3.28B.
- Data management, processing: $31.6B × 2.00%
- Pharmaceutical: $10.6B × 10.00%
- Accounting/finance: $2.59B × 25.00%
- Legal tech: $3.79B × 25.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on extracting and organizing knowledge. Learning on graphs is covered under graph ML.
Organizations
Selected examples- ABBYY Product
Document AI.
- Hyperscience Product
Document processing.
- Instabase Product
Enterprise document AI.
More organizations (4)
- Google Cloud Tools & infrastructure
Document AI extracts structured fields and entities from documents.
- Kensho Product
Financial AI.
- Neo4j Product
Graph data science and knowledge graphs.
- Rossum Product
Transactional document AI.
Dialogue systems
Conversation management, grounded dialogue, question answering, and language interfaces.
- Authors · scenario
- 3.0k–8.2k
- Investment · scenario
- $1.08B–$4.32B
Sources and methodology
Publication activity
- 2019
- 569
- 2020
- 775
- 2021
- 859
- 2022
- 978
- 2023
- 1,375
- 2024
- 1,569
- 2025
- 2,586
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("dialogue system" OR "dialog system" OR "conversational AI" OR "conversational agent" OR "task oriented dialogue" OR "conversational question answering")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.0k publishing authors. Paper count × 3.19 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,978 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,586 papers (38.7%). 14.7% of authorship records lack an Author ID; 12.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.19; 95% bootstrap interval 2.97–3.40. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.16B.
- AI agents: $8.02B × 10.00%
- Retail: $4.10B × 10.00%
- Medical and healthcare: $11.8B × 5.00%
- Ed tech: $1.44B × 25.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Agent research covers action execution; human–AI interaction covers how people use these systems.
Organizations
Selected examples- Sierra Product
Customer-service agents.
- Intercom Product
AI customer support.
- Convai Product
Embodied conversational characters.
More organizations (7)
- Anthropic Product
Claude provides conversational assistance for analysis and complex tasks.
- Decagon Product
Customer-service agents.
- ElevenLabs Product
Speech and voice agents.
- Google DeepMind Models
Gemini models support multimodal dialogue and question answering.
- OpenAI Product
ChatGPT supports contextual conversations and question answering.
- Replika Product
Conversational companions.
- Speak Product
AI language tutor.
Text understanding and generation
Semantics, syntax, discourse, classification, sentiment, summarization, and task-specific text generation.
- Authors · scenario
- 6.6k–33k
- Investment · scenario
- $1.02B–$4.10B
Sources and methodology
Publication activity
- 2019
- 3,369
- 2020
- 4,518
- 2021
- 5,209
- 2022
- 5,650
- 2023
- 7,383
- 2024
- 9,293
- 2025
- 11,797
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("natural language understanding" OR "text classification" OR "sentiment analysis" OR "text summarization" OR "semantic parsing" OR "discourse parsing" OR "natural language generation")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 11k publishing authors. Paper count × 2.80 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,757 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 11,797 papers (8.5%). 12.0% of authorship records lack an Author ID; 15.7% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 2.80; 95% bootstrap interval 2.68–2.92. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.05B.
- AI infrastructure/models/research/governance: $143.2B × 0.80%
- Legal tech: $3.79B × 20.00%
- Ed tech: $1.44B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Groups language tasks rather than the architectures used to solve them.
Organizations
Selected examples- Harvey Product
Legal AI.
- Writer Product
Enterprise AI agents.
- Superhuman / Grammarly Product
AI writing assistance through the Grammarly product.
More organizations (4)
- Anthropic Product
Claude supports writing, editing and text analysis.
- DeepL Product
Machine translation.
- Google DeepMind Models
Gemini models generate and analyze natural-language text.
- OpenAI Product
ChatGPT supports drafting, rewriting and summarizing text.