Speech and audio AI Sections

Speech and audio AI

Recognizing and generating speech, working with music and understanding the wider acoustic environment.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Speech recognition and understanding

Automatic speech recognition, speech translation, spoken-language understanding, and low-resource speech.

ASR Speech-to-text
1,348 papers
1.2× vs. 2023
Authors · scenario
3.3k–5.3k
Investment · scenario
$438M–$1.75B
Sources and methodology

Publication activity

2019
520
2020
748
2021
933
2022
1,042
2023
1,150
2024
1,197
2025
1,348

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("automatic speech recognition" OR "speech translation" OR "spoken language understanding" OR "end to end speech recognition")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.3k publishing authors. Paper count × 3.90 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,346 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,348 papers (74.2%). 12.1% of authorship records lack an Author ID; 11.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.90; 95% bootstrap interval 3.66–4.17. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $876M.

  • Internet of things: $14.6B × 5.00%
  • Ed tech: $1.44B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Voice generation and non-speech audio analysis are separate areas.

Organizations

Selected examples
  • OpenAI Models

    Speech models transcribe audio into text.

  • Google Cloud Tools & infrastructure

    Chirp and Speech-to-Text APIs provide multilingual speech recognition.

  • Deepgram Product

    Voice AI infrastructure.

More organizations (6)

Speech generation and voice AI

Text-to-speech, voice conversion, expressive speech, and speech-to-speech models.

TTS Text-to-speech Speech-to-speech
2,193 papers
1.5× vs. 2023
Authors · scenario
3.3k–7.9k
Investment · scenario
$282M–$1.13B
Sources and methodology

Publication activity

2019
831
2020
1,042
2021
1,133
2022
1,240
2023
1,417
2024
1,644
2025
2,193

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("text to speech" OR "speech synthesis" OR "voice conversion" OR "speech to speech")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.3k publishing authors. Paper count × 3.58 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,304 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,193 papers (45.6%). 12.6% of authorship records lack an Author ID; 9.8% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.58; 95% bootstrap interval 3.37–3.83. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $563M.

  • Creative, music, video content: $4.19B × 10.00%
  • Ed tech: $1.44B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Focuses on speech; music generation is listed separately.

Organizations

Selected examples
  • ElevenLabs Product

    Speech and voice agents.

  • OpenAI Models

    Text-to-speech models generate expressive spoken audio.

  • Cartesia Product

    Real-time audio models.

More organizations (3)

Music AI

Music generation, transcription, structural analysis, and music retrieval.

Music generation Music retrieval
525 papers
1.7× vs. 2023
Authors · observed
1,311
Investment · scenario
$392M–$1.57B
Sources and methodology

Publication activity

2019
144
2020
208
2021
225
2022
256
2023
317
2024
409
2025
525

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("music generation" OR "automatic music transcription" OR "music information retrieval" OR "music understanding" OR "music source separation")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

1,311 distinct OpenAlex Author IDs across all 525 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 525 of 525 papers (100.0%). 17.1% of authorship records lack an Author ID; 8.0% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $785M.

  • Creative, music, video content: $4.19B × 10.00%
  • Entertainment: $3.66B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Groups tasks that work with musical signals, across both analysis and generation.

Organizations

Selected examples
  • Google DeepMind Models

    Lyria models generate music and audio.

  • Suno Product

    Music generation.

  • Udio Product

    Music generation.

More organizations (3)

Audio perception and processing

Sound events, speaker recognition, diarization, source separation, enhancement, and spatial audio.

Diarization Audio separation Spatial audio
956 papers
0.9× vs. 2023
Authors · observed
2,958
Investment · scenario
$732M–$2.93B
Sources and methodology

Publication activity

2019
628
2020
786
2021
821
2022
901
2023
1,010
2024
1,044
2025
956

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("sound event detection" OR "audio classification" OR "speaker recognition" OR "speaker diarization" OR "speech enhancement" OR "audio source separation" OR "acoustic scene classification")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

2,958 distinct OpenAlex Author IDs across all 956 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 956 of 956 papers (100.0%). 10.2% of authorship records lack an Author ID; 19.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.46B.

  • Internet of things: $14.6B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Covers audio tasks beyond speech recognition and synthesis.

Organizations

Selected examples
  • Picovoice Tools & infrastructure

    Koala, Eagle and Falcon provide noise suppression, speaker recognition and diarization.

  • Krisp Product

    Audio enhancement.

  • Sensory Product

    Sound ID detects acoustic events such as alarms, glass breaking and doorbells.

More organizations (2)