Generative AI and world models
Creating images, video, 3D scenes and motion, and learning models of interactive environments.
By Dave Savostyanov. 2025 data · Sources: 26 Sep 2026
2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.
Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.
Image generation and editing
Text-to-image, image-to-image, editing, and controlled synthesis using diffusion, flow, or autoregressive models.
- Authors · scenario
- 3.6k–11k
- Investment · scenario
- $953M–$3.81B
Sources and methodology
Publication activity
- 2019
- 241
- 2020
- 315
- 2021
- 439
- 2022
- 648
- 2023
- 1,768
- 2024
- 2,812
- 2025
- 2,989
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("text to image" OR "image generation" OR "image editing") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "diffusion model" OR "flow matching")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 3.8k publishing authors. Paper count × 3.78 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,640 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,989 papers (33.5%). 15.0% of authorship records lack an Author ID; 14.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 3.78; 95% bootstrap interval 3.60–3.95. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.91B.
- Creative, music, video content: $4.19B × 28.00%
- Entertainment: $3.66B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Grouped by the output task; the underlying methods also apply to other modalities.
Organizations
Selected examples- OpenAI Models
Native image generation and editing in multimodal models.
- Google DeepMind Models
Image generation and editing models in the Gemini model family.
- Black Forest Labs Models
Image-generation models.
More organizations (9)
- Adobe / Firefly Product
Firefly generates images from text prompts.
- Ideogram Product
Text-to-image generation and editing workflows.
- Kuaishou / Kling AI Product
AI video generation.
- Lightricks Product
Generative creative tools.
- Luma AI Product
Visual generative models.
- Midjourney Models
Image and video generation models.
- Photoroom Product
Product imagery and editing.
- Recraft Product
Generative design.
- Stability AI Models
Generative media models.
Video generation and editing
Text- and image-to-video, temporal consistency, scene control, and video editing.
- Authors · observed
- 4,159
- Investment · scenario
- $1.04B–$4.15B
Sources and methodology
Publication activity
- 2019
- 39
- 2020
- 58
- 2021
- 72
- 2022
- 106
- 2023
- 283
- 2024
- 696
- 2025
- 954
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("text to video" OR "video generation" OR "video diffusion" OR "video editing") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "diffusion model" OR "generative model")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
4,159 distinct OpenAlex Author IDs across all 954 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 954 of 954 papers (100.0%). 15.9% of authorship records lack an Author ID; 10.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.07B.
- Creative, music, video content: $4.19B × 32.00%
- Entertainment: $3.66B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Action-conditioned interactive models also appear under world models.
Organizations
Selected examples- Google DeepMind Models
Veo generates video with synchronized audio.
- Runway Models
Video and world models.
- Luma AI Product
Visual generative models.
More organizations (10)
- Adobe / Firefly Product
Adobe Firefly generates video from text prompts and reference images.
- ByteDance Seed Models
Multimodal model research.
- Decart Models
Real-time generative video.
- HeyGen Product
Talking avatars and video translation.
- Kuaishou / Kling AI Product
AI video generation.
- Lightricks Product
Generative creative tools.
- Midjourney Models
Image and video generation models.
- Mirage / Captions Product
AI video models, avatars and editing tools.
- Pika Product
AI video generation.
- Synthesia Product
AI avatars and business video.
3D generation and neural rendering
Text- and image-to-3D, scene representations, NeRF, Gaussian splatting, and novel-view synthesis.
- Authors · scenario
- 4.2k–12k
- Investment · scenario
- $392M–$1.57B
Sources and methodology
Publication activity
- 2019
- 43
- 2020
- 85
- 2021
- 237
- 2022
- 455
- 2023
- 1,148
- 2024
- 2,251
- 2025
- 2,641
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("neural radiance field" OR "Gaussian splatting" OR "text to 3D" OR "3D generation" OR "neural rendering" OR "novel view synthesis")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
Base scenario: 4.2k publishing authors. Paper count × 4.68 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,155 distinct Author IDs already observed in the sample.
The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.
Random sample: 1,000 of 2,641 papers (37.9%). 13.7% of authorship records lack an Author ID; 13.3% of retrieved works lack an abstract. Retrieved 2026-09-26.
Mean known authors per paper: 4.68; 95% bootstrap interval 4.45–4.95. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $785M.
- Creative, music, video content: $4.19B × 10.00%
- Entertainment: $3.66B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Overlaps with geometric vision, especially when reconstructing and rendering the same scene.
Organizations
Selected examples- Microsoft Research
TRELLIS generates 3D assets from text and images.
- Meshy Product
Generative 3D.
- Tripo AI Product
Generative 3D.
More organizations (5)
- Kaedim Product
AI-assisted 3D assets.
- Sloyd Product
3D asset generation.
- SpAItial Models
Spatial foundation models.
- Tencent / Hunyuan Models
Hunyuan3D generates textured 3D assets and provides model weights and code.
- World Labs Models
Spatial intelligence and world models.
Motion generation and avatars
Text-to-motion, speech-driven gestures, talking heads, retargeting, and digital characters.
- Authors · observed
- 2,132
- Investment · scenario
- $576M–$2.30B
Sources and methodology
Publication activity
- 2019
- 105
- 2020
- 103
- 2021
- 125
- 2022
- 174
- 2023
- 301
- 2024
- 463
- 2025
- 553
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("motion generation" OR "text to motion" OR "talking head generation" OR "gesture generation" OR "motion retargeting" OR "animatable avatar")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
2,132 distinct OpenAlex Author IDs across all 553 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 553 of 553 papers (100.0%). 14.3% of authorship records lack an Author ID; 11.6% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $1.15B.
- Creative, music, video content: $4.19B × 10.00%
- Entertainment: $3.66B × 20.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Creates movement and appearance. Pose estimation and human recognition are covered under vision.
Organizations
Selected examples- DeepMotion Product
SayMotion generates 3D character animation from text.
- Motorica Product
Generative character animation.
- Kinetix Product
Motion and character animation.
More organizations (8)
- Autodesk / Wonder Dynamics Product
AI-assisted character animation and visual-effects workflows.
Part of Autodesk - Brahma / Metaphysic Product
AI-assisted digital-human content creation.
Part of Brahma / DNEG Group - D-ID Product
Digital avatars.
- HeyGen Product
Talking avatars and video translation.
- Mirage / Captions Product
AI video models, avatars and editing tools.
- Move AI Product
Markerless motion capture.
- Plask Product
AI motion capture.
- Synthesia Product
AI avatars and business video.
World models and simulation
Learned dynamics, action-conditioned generation, interactive environments, and simulation.
- Authors · observed
- 1,556
- Investment · scenario
- $1.41B–$5.63B
Sources and methodology
Publication activity
- 2019
- 23
- 2020
- 49
- 2021
- 77
- 2022
- 83
- 2023
- 150
- 2024
- 231
- 2025
- 510
Matches include papers applying these methods. Search precision and recall have not been systematically measured.
Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.
("world model" OR "action conditioned generation" OR "interactive simulation") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "reinforcement learning")Retrieved 2026-09-26.
OpenAlex query resultsPublishing authors
1,556 distinct OpenAlex Author IDs across all 510 matching papers. This is an observed count within the query result; no publication-rate assumption is used.
Full query result: 510 of 510 papers (100.0%). 17.2% of authorship records lack an Author ID; 10.2% of retrieved works lack an abstract. Retrieved 2026-09-26.
Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.
OpenAlex sample queryPrivate investment
An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.
Base case: $2.81B.
- AI infrastructure/models/research/governance: $143.2B × 1.00%
- Robotics: $7.84B × 10.00%
- Energy management: $4.64B × 5.00%
- Entertainment: $3.66B × 10.00%
Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.
Stanford AI Index 2026 / QuidScope and related research
Model-based reinforcement learning uses these models for decision-making.
Organizations
Selected examples- World Labs Models
Spatial intelligence and world models.
- Google DeepMind Research
Models, world models and robotics.
- Odyssey Models
Interactive world models.
More organizations (7)
- Decart Models
Real-time generative video.
- NVIDIA Research
Physical AI and model research.
- Runway Models
Video and world models.
- SpAItial Models
Spatial foundation models.
- Tencent / Hunyuan Research
HY-World reconstructs, generates and simulates 3D environments.
- Waabi Product
Autonomous trucking and simulation.
- Wayve Research
Research on driving world models, 3D perception and reconstruction.