Search and recommender systems
Finding and ranking information, recommending relevant content and allocating advertising opportunities.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Information retrieval and RAG
Sparse and dense retrieval, embeddings, reranking, learning-to-rank, indexing, and retrieval-augmented generation.
- Authors · scenario
- 5.6k–28k
- Investment · scenario
- $2.46B–$9.85B
Sources and methodology
Publication activity
- 2019
- 1,696
- 2020
- 1,927
- 2021
- 1,874
- 2022
- 1,888
- 2023
- 2,265
- 2024
- 4,234
- 2025
- 7,700
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("information retrieval" OR "learning to rank" OR "dense retrieval" OR "neural retrieval" OR "retrieval augmented generation" OR "retrieval augmented language model")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 9.4k publishing authors. Paper count × 3.64 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,564 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 7,700 papers (13.0%). 12.1% of authorship records lack an Author ID; 9.9% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.64; 95% bootstrap interval 3.45–3.87. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $4.92B.
- Data management, processing: $31.6B × 8.00%
- Marketing, digital ads: $2.51B × 10.00%
- Retail: $4.10B × 20.00%
- Legal tech: $3.79B × 35.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
RAG is one part of the broader information-retrieval field.
Organizations
Selected examples- Cohere Models
Embed and Rerank models support semantic search and relevance ranking.
- Pinecone Tools & infrastructure
Vector search.
- Weaviate Tools & infrastructure
Vector database.
More organizations (12)
- Algolia Product
Search and recommendations.
- AlphaSense Product
Financial and market research.
- Coveo Product
Enterprise search and recommendations.
- Elastic Tools & infrastructure
Search and retrieval infrastructure.
- Glean Product
Enterprise search and agents.
- Perplexity Tools & infrastructure
Search and embedding APIs provide web retrieval, ranking and semantic search.
- Qdrant Tools & infrastructure
Vector database.
- Sana Product
AI learning and knowledge.
Part of Workday - Sourcegraph Product
Code intelligence.
- TwelveLabs Models
Marengo provides multimodal embeddings and semantic search over video moments.
- Vespa Tools & infrastructure
Search and recommendation serving.
- Zilliz Tools & infrastructure
Vector database and Milvus.
Recommender systems
Candidate retrieval, sequential and generative recommendation, and user-interest modeling.
- Authors · scenario
- 4.5k–22k
- Investment · scenario
- $1.46B–$5.84B
Sources and methodology
Publication activity
- 2019
- 2,637
- 2020
- 3,119
- 2021
- 3,493
- 2022
- 3,876
- 2023
- 4,603
- 2024
- 5,826
- 2025
- 7,465
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("recommender system" OR "recommendation system" OR "sequential recommendation" OR "generative recommendation" OR "collaborative filtering")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 7.5k publishing authors. Paper count × 3.01 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 2,912 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 7,465 papers (13.4%). 14.9% of authorship records lack an Author ID; 16.1% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.01; 95% bootstrap interval 2.87–3.15. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.92B.
- Marketing, digital ads: $2.51B × 20.00%
- Retail: $4.10B × 50.00%
- Entertainment: $3.66B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Advertising auctions and bidding are listed separately, although the methods overlap.
Organizations
Selected examplesAds ranking and auction design
Click and conversion prediction, attribution, bidding, auctions, and ad-delivery optimization.
- Authors · observed
- 443
- Investment · scenario
- $816M–$3.26B
Sources and methodology
Publication activity
- 2019
- 63
- 2020
- 102
- 2021
- 93
- 2022
- 132
- 2023
- 116
- 2024
- 129
- 2025
- 160
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("click through rate prediction" OR "conversion rate prediction" OR "computational advertising" OR "real time bidding" OR "ad auction" OR "advertising attribution")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
443 distinct OpenAlex Author IDs across all 160 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 160 of 160 papers (100.0%). 6.8% of authorship records lack an Author ID; 21.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.63B.
- Marketing, digital ads: $2.51B × 65.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Draws on ranking, causal inference, and mechanism design.
Organizations
Selected examples- Google Ads Product
Performance Max applies Google AI to bidding and ad delivery.
- Meta Research
Andromeda applies machine learning to personalized ad retrieval.
- Moloco Product
ML for advertising and commerce.
More organizations (3)
- AppLovin Product
Advertising optimization.
- Criteo Product
Commerce advertising.
- The Trade Desk Product
Programmatic advertising.