Computer vision
Understanding images, video, documents, people and three-dimensional scenes.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Image recognition, detection and segmentation
Image classification, object detection, segmentation, open-vocabulary recognition, and visual retrieval.
- Authors · scenario
- 19k–97k
- Investment · scenario
- $2.77B–$11.1B
Sources and methodology
Publication activity
- 2019
- 6,340
- 2020
- 9,014
- 2021
- 11,847
- 2022
- 14,600
- 2023
- 17,851
- 2024
- 21,507
- 2025
- 24,491
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("image classification" OR "object detection" OR "semantic segmentation" OR "instance segmentation" OR "open vocabulary recognition")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 32k publishing authors. Paper count × 3.96 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,911 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 24,491 papers (4.1%). 11.8% of authorship records lack an Author ID; 24.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.96; 95% bootstrap interval 3.79–4.15. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $5.54B.
- Retail: $4.10B × 10.00%
- Internet of things: $14.6B × 15.00%
- Medical and healthcare: $11.8B × 25.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Medical and industrial uses also appear in the applications view.
Organizations
Selected examplesVideo understanding and tracking
Action and event recognition, temporal localization, tracking, optical flow, and video question answering.
- Authors · scenario
- 4.0k–20k
- Investment · scenario
- $198M–$794M
Sources and methodology
Publication activity
- 2019
- 2,192
- 2020
- 2,498
- 2021
- 2,885
- 2022
- 3,252
- 2023
- 3,754
- 2024
- 4,317
- 2025
- 4,819
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("action recognition" OR "video understanding" OR "video question answering" OR "temporal action localization" OR "object tracking" OR "optical flow")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 6.7k publishing authors. Paper count × 4.18 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,038 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,819 papers (20.8%). 12.6% of authorship records lack an Author ID; 21.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.18; 95% bootstrap interval 4.02–4.34. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $397M.
- Autonomous vehicles: $7.94B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Video generation is a separate area. Vision-language models can support these tasks.
Organizations
Selected examples- TwelveLabs Models
Marengo and Pegasus model visual, audio and temporal information in video.
- Google DeepMind Models
Gemini models process video for temporal understanding and question answering.
- Meta Models
SAM 2 extends promptable segmentation and object tracking to video.
3D vision and spatial perception
Depth, stereo, structure from motion, reconstruction, pose estimation, and scene understanding.
- Authors · scenario
- 6.3k–32k
- Investment · scenario
- $198M–$794M
Sources and methodology
Publication activity
- 2019
- 2,636
- 2020
- 3,045
- 2021
- 3,361
- 2022
- 3,669
- 2023
- 4,372
- 2024
- 5,220
- 2025
- 6,743
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("depth estimation" OR "stereo matching" OR "structure from motion" OR "3D reconstruction" OR "point cloud understanding" OR "scene understanding")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 11k publishing authors. Paper count × 4.70 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,532 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 6,743 papers (14.8%). 10.5% of authorship records lack an Author ID; 22.0% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.70; 95% bootstrap interval 4.51–4.88. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $397M.
- Autonomous vehicles: $7.94B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Reconstructs observed scenes. Generating new 3D content and robot navigation are related areas.
Organizations
Selected examples- NVIDIA Models
FoundationPose and FoundationStereo estimate object pose and stereo depth.
- Wayve Research
Research on driving world models, 3D perception and reconstruction.
- Skydio Product
Autonomous drones.
More organizations (2)
- Exyn Technologies Product
Autonomous aerial mapping.
- Tencent / Hunyuan Research
HY-World reconstructs, generates and simulates 3D environments.
Image restoration
Denoising, deblurring, super-resolution, inverse imaging, and physics-based vision.
- Authors · scenario
- 3.8k–5.7k
- Investment · scenario
- $588M–$2.35B
Sources and methodology
Publication activity
- 2019
- 521
- 2020
- 748
- 2021
- 919
- 2022
- 1,070
- 2023
- 1,235
- 2024
- 1,297
- 2025
- 1,386
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("image restoration" OR "image denoising" OR "image deblurring" OR "image super resolution" OR "computational imaging") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.8k publishing authors. Paper count × 4.13 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,810 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,386 papers (72.2%). 10.5% of authorship records lack an Author ID; 27.7% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.13; 95% bootstrap interval 3.97–4.32. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.18B.
- Medical and healthcare: $11.8B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Recovers a signal from measurements. Unconstrained image generation is listed separately.
Organizations
Selected examples- Topaz Labs Product
Image and video enhancement.
- DxO Product
DeepPRIME uses neural networks for RAW-image denoising and demosaicing.
- Adobe / Firefly Product
Camera Raw provides AI denoising and Super Resolution image enhancement.
More organizations (1)
- Photoroom Product
Product imagery and editing.
OCR and document understanding
Text recognition, document layout, tables, formulas, document question answering, and structured extraction.
- Authors · scenario
- 3.1k–4.6k
- Investment · scenario
- $1.00B–$4.00B
Sources and methodology
Publication activity
- 2019
- 421
- 2020
- 474
- 2021
- 548
- 2022
- 562
- 2023
- 749
- 2024
- 865
- 2025
- 1,397
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("optical character recognition" OR "document understanding" OR "document layout analysis" OR "document question answering" OR "table recognition" OR "scene text recognition")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.1k publishing authors. Paper count × 3.29 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,131 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 1,397 papers (71.6%). 14.9% of authorship records lack an Author ID; 16.0% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.29; 95% bootstrap interval 3.07–3.57. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.00B.
- Medical and healthcare: $11.8B × 5.00%
- Accounting/finance: $2.59B × 40.00%
- Legal tech: $3.79B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Combines vision and language; vision-language models are one implementation approach.
Organizations
Selected examples- Google Cloud Tools & infrastructure
Document AI recognizes text, layout and tables in documents.
- Microsoft Tools & infrastructure
Document Intelligence recognizes printed and handwritten text and document layout.
- ABBYY Product
Document AI.
More organizations (7)
- Hyperscience Product
Document processing.
- Instabase Product
Enterprise document AI.
- LandingAI Product
Document extraction and document automation.
- LlamaIndex Product
Document parsing and agents over enterprise information.
- Rossum Product
Transactional document AI.
- Upstage Models
Document AI and language models.
- V7 Product
Document-centered workflow automation.
Human perception and biometrics
Face, body, hand and pose analysis, gesture recognition, motion analysis, and identification.
- Authors · scenario
- 3.5k–17k
- Investment · scenario
- $196M–$784M
Sources and methodology
Publication activity
- 2019
- 2,823
- 2020
- 3,198
- 2021
- 3,439
- 2022
- 3,657
- 2023
- 4,039
- 2024
- 4,181
- 2025
- 4,626
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("human pose estimation" OR "face recognition" OR "gesture recognition" OR "human mesh recovery" OR "hand pose estimation" OR "person re identification")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 5.6k publishing authors. Paper count × 3.62 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,517 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 4,626 papers (21.6%). 10.9% of authorship records lack an Author ID; 17.5% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.62; 95% bootstrap interval 3.48–3.77. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $392M.
- Robotics: $7.84B × 5.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Analyzes observed people. Creating motion and avatars is covered under generation.
Organizations
Selected examples- Move AI Product
Markerless motion capture.
- Innovatrics Product
Biometrics.
- Plask Product
AI motion capture.
More organizations (3)
- DeepMotion Product
Animate 3D extracts body motion from video for animation.
- IDEMIA Product
Biometric identity systems.
- Oosto Product
Visual intelligence.