Speech and audio AI
Recognizing and generating speech, working with music and understanding the wider acoustic environment.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Speech recognition and understanding
Automatic speech recognition, speech translation, spoken-language understanding, and low-resource speech.
- Authors · scenario
- 3.3k–5.3k
- Investment · scenario
- $438M–$1.75B
Sources and methodology
Publication activity
- 2019
- 520
- 2020
- 748
- 2021
- 933
- 2022
- 1,042
- 2023
- 1,150
- 2024
- 1,197
- 2025
- 1,348
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("automatic speech recognition" OR "speech translation" OR "spoken language understanding" OR "end to end speech recognition")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.3k publishing authors. Paper count × 3.90 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,346 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,348 papers (74.2%). 12.1% of authorship records lack an Author ID; 11.7% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.90; 95% bootstrap interval 3.66–4.17. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $876M.
- Internet of things: $14.6B × 5.00%
- Ed tech: $1.44B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Voice generation and non-speech audio analysis are separate areas.
Organizations
Selected examples- OpenAI Models
Speech models transcribe audio into text.
- Google Cloud Tools & infrastructure
Chirp and Speech-to-Text APIs provide multilingual speech recognition.
- Deepgram Product
Voice AI infrastructure.
More organizations (6)
- Abridge Product
Clinical documentation.
- AssemblyAI Product
Speech recognition.
- ELSA Product
AI English speaking tutor.
- Inworld AI Product
Speech models and inference APIs.
- Nabla Product
Clinical assistant.
- Speechmatics Product
Speech recognition.
Speech generation and voice AI
Text-to-speech, voice conversion, expressive speech, and speech-to-speech models.
- Authors · scenario
- 3.3k–7.9k
- Investment · scenario
- $282M–$1.13B
Sources and methodology
Publication activity
- 2019
- 831
- 2020
- 1,042
- 2021
- 1,133
- 2022
- 1,240
- 2023
- 1,417
- 2024
- 1,644
- 2025
- 2,193
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("text to speech" OR "speech synthesis" OR "voice conversion" OR "speech to speech")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.3k publishing authors. Paper count × 3.58 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,304 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,193 papers (45.6%). 12.6% of authorship records lack an Author ID; 9.8% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.58; 95% bootstrap interval 3.37–3.83. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $563M.
- Creative, music, video content: $4.19B × 10.00%
- Ed tech: $1.44B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Focuses on speech; music generation is listed separately.
Organizations
Selected examples- ElevenLabs Product
Speech and voice agents.
- OpenAI Models
Text-to-speech models generate expressive spoken audio.
- Cartesia Product
Real-time audio models.
More organizations (3)
- ByteDance Seed Models
Multimodal model research.
- Deepgram Product
Voice AI infrastructure.
- Inworld AI Product
Speech models and inference APIs.
Music AI
Music generation, transcription, structural analysis, and music retrieval.
- Authors · observed
- 1,311
- Investment · scenario
- $392M–$1.57B
Sources and methodology
Publication activity
- 2019
- 144
- 2020
- 208
- 2021
- 225
- 2022
- 256
- 2023
- 317
- 2024
- 409
- 2025
- 525
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("music generation" OR "automatic music transcription" OR "music information retrieval" OR "music understanding" OR "music source separation")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
1,311 distinct OpenAlex Author IDs across all 525 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 525 of 525 papers (100.0%). 17.1% of authorship records lack an Author ID; 8.0% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $785M.
- Creative, music, video content: $4.19B × 10.00%
- Entertainment: $3.66B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Groups tasks that work with musical signals, across both analysis and generation.
Organizations
Selected examples- Google DeepMind Models
Lyria models generate music and audio.
- Suno Product
Music generation.
- Udio Product
Music generation.
More organizations (3)
- AudioShake Product
Audio source separation.
- Boomy Product
AI music.
- SOUNDRAW Product
AI music.
Audio perception and processing
Sound events, speaker recognition, diarization, source separation, enhancement, and spatial audio.
- Authors · observed
- 2,958
- Investment · scenario
- $732M–$2.93B
Sources and methodology
Publication activity
- 2019
- 628
- 2020
- 786
- 2021
- 821
- 2022
- 901
- 2023
- 1,010
- 2024
- 1,044
- 2025
- 956
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("sound event detection" OR "audio classification" OR "speaker recognition" OR "speaker diarization" OR "speech enhancement" OR "audio source separation" OR "acoustic scene classification")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
2,958 distinct OpenAlex Author IDs across all 956 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 956 of 956 papers (100.0%). 10.2% of authorship records lack an Author ID; 19.7% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.46B.
- Internet of things: $14.6B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Covers audio tasks beyond speech recognition and synthesis.
Organizations
Selected examples- Picovoice Tools & infrastructure
Koala, Eagle and Falcon provide noise suppression, speaker recognition and diarization.
- Krisp Product
Audio enhancement.
- Sensory Product
Sound ID detects acoustic events such as alarms, glass breaking and doorbells.
More organizations (2)
- ai-coustics Product
Speech enhancement.
- AudioShake Product
Audio source separation.