Computer vision Sections

Computer vision

Understanding images, video, documents, people and three-dimensional scenes.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Image recognition, detection and segmentation

Image classification, object detection, segmentation, open-vocabulary recognition, and visual retrieval.

Object detection Segmentation Open-vocabulary vision
24,491 papers
1.4× vs. 2023
Authors · scenario
19k–97k
Investment · scenario
$2.77B–$11.1B
Sources and methodology

Publication activity

2019
6,340
2020
9,014
2021
11,847
2022
14,600
2023
17,851
2024
21,507
2025
24,491

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("image classification" OR "object detection" OR "semantic segmentation" OR "instance segmentation" OR "open vocabulary recognition")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 32k publishing authors. Paper count × 3.96 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,911 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 24,491 papers (4.1%). 11.8% of authorship records lack an Author ID; 24.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.96; 95% bootstrap interval 3.79–4.15. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $5.54B.

  • Retail: $4.10B × 10.00%
  • Internet of things: $14.6B × 15.00%
  • Medical and healthcare: $11.8B × 25.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Medical and industrial uses also appear in the applications view.

Organizations

Selected examples
  • Meta Models

    Segment Anything models perform promptable image segmentation.

  • Roboflow Product

    Computer vision development.

  • Cognex Product

    Deep-learning tools for industrial image inspection and recognition.

More organizations (5)
  • Mobileye Product

    Driving perception and autonomy.

  • OpenAI Research

    CLIP supports open-vocabulary image classification using text descriptions.

  • Paige Product

    AI pathology.

    Part of Tempus
  • PathAI Product

    AI pathology.

  • Qure.ai Product

    Medical imaging AI.

Video understanding and tracking

Action and event recognition, temporal localization, tracking, optical flow, and video question answering.

Video understanding Tracking
4,819 papers
1.3× vs. 2023
Authors · scenario
4.0k–20k
Investment · scenario
$198M–$794M
Sources and methodology

Publication activity

2019
2,192
2020
2,498
2021
2,885
2022
3,252
2023
3,754
2024
4,317
2025
4,819

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("action recognition" OR "video understanding" OR "video question answering" OR "temporal action localization" OR "object tracking" OR "optical flow")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 6.7k publishing authors. Paper count × 4.18 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,038 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,819 papers (20.8%). 12.6% of authorship records lack an Author ID; 21.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.18; 95% bootstrap interval 4.02–4.34. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $397M.

  • Autonomous vehicles: $7.94B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Video generation is a separate area. Vision-language models can support these tasks.

Organizations

Selected examples
  • TwelveLabs Models

    Marengo and Pegasus model visual, audio and temporal information in video.

  • Google DeepMind Models

    Gemini models process video for temporal understanding and question answering.

  • Meta Models

    SAM 2 extends promptable segmentation and object tracking to video.

More organizations (2)
  • Oosto Product

    Visual intelligence.

  • Reka AI Models

    Multimodal models.

3D vision and spatial perception

Depth, stereo, structure from motion, reconstruction, pose estimation, and scene understanding.

Spatial AI 3D reconstruction Depth estimation
6,743 papers
1.5× vs. 2023
Authors · scenario
6.3k–32k
Investment · scenario
$198M–$794M
Sources and methodology

Publication activity

2019
2,636
2020
3,045
2021
3,361
2022
3,669
2023
4,372
2024
5,220
2025
6,743

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("depth estimation" OR "stereo matching" OR "structure from motion" OR "3D reconstruction" OR "point cloud understanding" OR "scene understanding")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 11k publishing authors. Paper count × 4.70 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,532 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 6,743 papers (14.8%). 10.5% of authorship records lack an Author ID; 22.0% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.70; 95% bootstrap interval 4.51–4.88. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $397M.

  • Autonomous vehicles: $7.94B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Reconstructs observed scenes. Generating new 3D content and robot navigation are related areas.

Organizations

Selected examples
  • NVIDIA Models

    FoundationPose and FoundationStereo estimate object pose and stereo depth.

  • Wayve Research

    Research on driving world models, 3D perception and reconstruction.

  • Skydio Product

    Autonomous drones.

More organizations (2)

Image restoration

Denoising, deblurring, super-resolution, inverse imaging, and physics-based vision.

Super-resolution Computational imaging
1,386 papers
1.1× vs. 2023
Authors · scenario
3.8k–5.7k
Investment · scenario
$588M–$2.35B
Sources and methodology

Publication activity

2019
521
2020
748
2021
919
2022
1,070
2023
1,235
2024
1,297
2025
1,386

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("image restoration" OR "image denoising" OR "image deblurring" OR "image super resolution" OR "computational imaging") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence"))

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.8k publishing authors. Paper count × 4.13 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,810 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,386 papers (72.2%). 10.5% of authorship records lack an Author ID; 27.7% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.13; 95% bootstrap interval 3.97–4.32. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.18B.

  • Medical and healthcare: $11.8B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Recovers a signal from measurements. Unconstrained image generation is listed separately.

Organizations

Selected examples
  • Topaz Labs Product

    Image and video enhancement.

  • DxO Product

    DeepPRIME uses neural networks for RAW-image denoising and demosaicing.

  • Adobe / Firefly Product

    Camera Raw provides AI denoising and Super Resolution image enhancement.

More organizations (1)
  • Photoroom Product

    Product imagery and editing.

OCR and document understanding

Text recognition, document layout, tables, formulas, document question answering, and structured extraction.

Document AI OCR
1,397 papers
1.9× vs. 2023
Authors · scenario
3.1k–4.6k
Investment · scenario
$1.00B–$4.00B
Sources and methodology

Publication activity

2019
421
2020
474
2021
548
2022
562
2023
749
2024
865
2025
1,397

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("optical character recognition" OR "document understanding" OR "document layout analysis" OR "document question answering" OR "table recognition" OR "scene text recognition")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.1k publishing authors. Paper count × 3.29 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,131 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 1,397 papers (71.6%). 14.9% of authorship records lack an Author ID; 16.0% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.29; 95% bootstrap interval 3.07–3.57. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.00B.

  • Medical and healthcare: $11.8B × 5.00%
  • Accounting/finance: $2.59B × 40.00%
  • Legal tech: $3.79B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Combines vision and language; vision-language models are one implementation approach.

Organizations

Selected examples
  • Google Cloud Tools & infrastructure

    Document AI recognizes text, layout and tables in documents.

  • Microsoft Tools & infrastructure

    Document Intelligence recognizes printed and handwritten text and document layout.

  • ABBYY Product

    Document AI.

More organizations (7)
  • Hyperscience Product

    Document processing.

  • Instabase Product

    Enterprise document AI.

  • LandingAI Product

    Document extraction and document automation.

  • LlamaIndex Product

    Document parsing and agents over enterprise information.

  • Rossum Product

    Transactional document AI.

  • Upstage Models

    Document AI and language models.

  • V7 Product

    Document-centered workflow automation.

Human perception and biometrics

Face, body, hand and pose analysis, gesture recognition, motion analysis, and identification.

Pose estimation Biometrics
4,626 papers
1.1× vs. 2023
Authors · scenario
3.5k–17k
Investment · scenario
$196M–$784M
Sources and methodology

Publication activity

2019
2,823
2020
3,198
2021
3,439
2022
3,657
2023
4,039
2024
4,181
2025
4,626

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("human pose estimation" OR "face recognition" OR "gesture recognition" OR "human mesh recovery" OR "hand pose estimation" OR "person re identification")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 5.6k publishing authors. Paper count × 3.62 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,517 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 4,626 papers (21.6%). 10.9% of authorship records lack an Author ID; 17.5% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.62; 95% bootstrap interval 3.48–3.77. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $392M.

  • Robotics: $7.84B × 5.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Analyzes observed people. Creating motion and avatars is covered under generation.

Organizations

Selected examples
More organizations (3)
  • DeepMotion Product

    Animate 3D extracts body motion from video for animation.

  • IDEMIA Product

    Biometric identity systems.

  • Oosto Product

    Visual intelligence.