Natural language processing Sections

Natural language processing

Working with language across translation, knowledge extraction, dialogue and text understanding.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Translation and multilingual NLP

Machine translation, cross-lingual transfer, language coverage, and low-resource methods.

Machine translation Multilingual LLMs
4,357 papers
1.6× vs. 2023
Authors · scenario
2.9k–13k
Investment · scenario
$681M–$2.72B
Sources and methodology

Publication activity

2019
1,574
2020
1,969
2021
2,203
2022
2,151
2023
2,688
2024
3,376
2025
4,357

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("machine translation" OR "multilingual language model" OR "cross lingual" OR "low resource language")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 4.5k publishing authors. Paper count × 3.07 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,912 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,357 papers (23.0%). 12.3% of authorship records lack an Author ID; 10.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.07; 95% bootstrap interval 2.89–3.27. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.36B.

  • AI infrastructure/models/research/governance: $143.2B × 0.80%
  • Ed tech: $1.44B × 15.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Language-specific quality and translation remain distinct tasks alongside general pre-training.

Organizations

Selected examples
  • DeepL Product

    Machine translation.

  • Google Cloud Tools & infrastructure

    Cloud Translation offers neural and language-model translation.

  • Cohere Models

    Rerank models rank relevant documents across languages.

More organizations (4)
  • Duolingo Product

    AI language learning.

  • Meta Models

    Llama 4 models support multilingual text understanding and generation.

  • Microsoft Tools & infrastructure

    Azure Translator provides neural and language-model text translation.

  • Unbabel Product

    Translation AI.

Information extraction and knowledge graphs

Entities, relations, events, knowledge graphs, ontologies, and neuro-symbolic knowledge representation.

Knowledge graphs Entity extraction
5,423 papers
1.9× vs. 2023
Authors · scenario
4.2k–21k
Investment · scenario
$1.64B–$6.57B
Sources and methodology

Publication activity

2019
868
2020
1,278
2021
1,526
2022
1,978
2023
2,793
2024
3,915
2025
5,423

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("information extraction" OR "named entity recognition" OR "relation extraction" OR "knowledge graph" OR "knowledge representation" OR "neurosymbolic") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR ("language model" OR "large language model" OR "vision language model") OR "natural language processing")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.9k publishing authors. Paper count × 3.84 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,724 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 5,423 papers (18.4%). 13.1% of authorship records lack an Author ID; 17.9% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.84; 95% bootstrap interval 3.60–4.09. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $3.28B.

  • Data management, processing: $31.6B × 2.00%
  • Pharmaceutical: $10.6B × 10.00%
  • Accounting/finance: $2.59B × 25.00%
  • Legal tech: $3.79B × 25.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on extracting and organizing knowledge. Learning on graphs is covered under graph ML.

Organizations

Selected examples
More organizations (4)
  • Google Cloud Tools & infrastructure

    Document AI extracts structured fields and entities from documents.

  • Kensho Product

    Financial AI.

  • Neo4j Product

    Graph data science and knowledge graphs.

  • Rossum Product

    Transactional document AI.

Dialogue systems

Conversation management, grounded dialogue, question answering, and language interfaces.

Conversational AI Dialogue systems
2,586 papers
1.9× vs. 2023
Authors · scenario
3.0k–8.2k
Investment · scenario
$1.08B–$4.32B
Sources and methodology

Publication activity

2019
569
2020
775
2021
859
2022
978
2023
1,375
2024
1,569
2025
2,586

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("dialogue system" OR "dialog system" OR "conversational AI" OR "conversational agent" OR "task oriented dialogue" OR "conversational question answering")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.0k publishing authors. Paper count × 3.19 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,978 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,586 papers (38.7%). 14.7% of authorship records lack an Author ID; 12.1% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.19; 95% bootstrap interval 2.97–3.40. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.16B.

  • AI agents: $8.02B × 10.00%
  • Retail: $4.10B × 10.00%
  • Medical and healthcare: $11.8B × 5.00%
  • Ed tech: $1.44B × 25.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Agent research covers action execution; human–AI interaction covers how people use these systems.

Organizations

Selected examples
  • Sierra Product

    Customer-service agents.

  • Intercom Product

    AI customer support.

  • Convai Product

    Embodied conversational characters.

More organizations (7)
  • Anthropic Product

    Claude provides conversational assistance for analysis and complex tasks.

  • Decagon Product

    Customer-service agents.

  • ElevenLabs Product

    Speech and voice agents.

  • Google DeepMind Models

    Gemini models support multimodal dialogue and question answering.

  • OpenAI Product

    ChatGPT supports contextual conversations and question answering.

  • Replika Product

    Conversational companions.

  • Speak Product

    AI language tutor.

Text understanding and generation

Semantics, syntax, discourse, classification, sentiment, summarization, and task-specific text generation.

NLP Text generation Summarization
11,797 papers
1.6× vs. 2023
Authors · scenario
6.6k–33k
Investment · scenario
$1.02B–$4.10B
Sources and methodology

Publication activity

2019
3,369
2020
4,518
2021
5,209
2022
5,650
2023
7,383
2024
9,293
2025
11,797

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("natural language understanding" OR "text classification" OR "sentiment analysis" OR "text summarization" OR "semantic parsing" OR "discourse parsing" OR "natural language generation")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 11k publishing authors. Paper count × 2.80 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,757 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 11,797 papers (8.5%). 12.0% of authorship records lack an Author ID; 15.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 2.80; 95% bootstrap interval 2.68–2.92. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.05B.

  • AI infrastructure/models/research/governance: $143.2B × 0.80%
  • Legal tech: $3.79B × 20.00%
  • Ed tech: $1.44B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Groups language tasks rather than the architectures used to solve them.

Organizations

Selected examples
More organizations (4)
  • Anthropic Product

    Claude supports writing, editing and text analysis.

  • DeepL Product

    Machine translation.

  • Google DeepMind Models

    Gemini models generate and analyze natural-language text.

  • OpenAI Product

    ChatGPT supports drafting, rewriting and summarizing text.