Generative AI and world models Sections

Generative AI and world models

Creating images, video, 3D scenes and motion, and learning models of interactive environments.

2025 figures. Papers are OpenAlex query matches; author ranges and investment allocations are scenarios. Author counts marked “observed” cover the full query result. Areas overlap and cannot be added together. About the data.

Bars compare publication counts within this page. Topic tags describe subtopics, not separately measured markets. Organization examples link to their sources; activity was checked in September 2026.

Image generation and editing

Text-to-image, image-to-image, editing, and controlled synthesis using diffusion, flow, or autoregressive models.

Text-to-image Diffusion models
2,989 papers
1.7× vs. 2023
Authors · scenario
3.6k–11k
Investment · scenario
$953M–$3.81B
Sources and methodology

Publication activity

2019
241
2020
315
2021
439
2022
648
2023
1,768
2024
2,812
2025
2,989

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("text to image" OR "image generation" OR "image editing") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "diffusion model" OR "flow matching")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 3.8k publishing authors. Paper count × 3.78 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 3,640 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,989 papers (33.5%). 15.0% of authorship records lack an Author ID; 14.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 3.78; 95% bootstrap interval 3.60–3.95. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.91B.

  • Creative, music, video content: $4.19B × 28.00%
  • Entertainment: $3.66B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Grouped by the output task; the underlying methods also apply to other modalities.

Organizations

Selected examples
  • OpenAI Models

    Native image generation and editing in multimodal models.

  • Google DeepMind Models

    Image generation and editing models in the Gemini model family.

  • Black Forest Labs Models

    Image-generation models.

More organizations (9)

Video generation and editing

Text- and image-to-video, temporal consistency, scene control, and video editing.

Text-to-video Video editing
954 papers
3.4× vs. 2023
Authors · observed
4,159
Investment · scenario
$1.04B–$4.15B
Sources and methodology

Publication activity

2019
39
2020
58
2021
72
2022
106
2023
283
2024
696
2025
954

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("text to video" OR "video generation" OR "video diffusion" OR "video editing") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "diffusion model" OR "generative model")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

4,159 distinct OpenAlex Author IDs across all 954 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 954 of 954 papers (100.0%). 15.9% of authorship records lack an Author ID; 10.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.07B.

  • Creative, music, video content: $4.19B × 32.00%
  • Entertainment: $3.66B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Action-conditioned interactive models also appear under world models.

Organizations

Selected examples
  • Google DeepMind Models

    Veo generates video with synchronized audio.

  • Runway Models

    Video and world models.

  • Luma AI Product

    Visual generative models.

More organizations (10)

3D generation and neural rendering

Text- and image-to-3D, scene representations, NeRF, Gaussian splatting, and novel-view synthesis.

3D generation Gaussian splatting NeRF
2,641 papers
2.3× vs. 2023
Authors · scenario
4.2k–12k
Investment · scenario
$392M–$1.57B
Sources and methodology

Publication activity

2019
43
2020
85
2021
237
2022
455
2023
1,148
2024
2,251
2025
2,641

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("neural radiance field" OR "Gaussian splatting" OR "text to 3D" OR "3D generation" OR "neural rendering" OR "novel view synthesis")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

Base scenario: 4.2k publishing authors. Paper count × 4.68 known authors per paper ÷ 3 papers per author per year. The range assumes 1–5 papers per author; the base and both ends of the range are at least the 4,155 distinct Author IDs already observed in the sample.

The publication rate is an assumption, not a measured rate for this field. The range is a scenario, not a confidence interval.

Random sample: 1,000 of 2,641 papers (37.9%). 13.7% of authorship records lack an Author ID; 13.3% of retrieved works lack an abstract. Retrieved 2026-09-26.

Mean known authors per paper: 4.68; 95% bootstrap interval 4.45–4.95. This describes sampling variability, not search accuracy or uncertainty in the number of researchers.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $785M.

  • Creative, music, video content: $4.19B × 10.00%
  • Entertainment: $3.66B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Overlaps with geometric vision, especially when reconstructing and rendering the same scene.

Organizations

Selected examples
  • Microsoft Research

    TRELLIS generates 3D assets from text and images.

  • Meshy Product

    Generative 3D.

  • Tripo AI Product

    Generative 3D.

More organizations (5)
  • Kaedim Product

    AI-assisted 3D assets.

  • Sloyd Product

    3D asset generation.

  • SpAItial Models

    Spatial foundation models.

  • Tencent / Hunyuan Models

    Hunyuan3D generates textured 3D assets and provides model weights and code.

  • World Labs Models

    Spatial intelligence and world models.

Motion generation and avatars

Text-to-motion, speech-driven gestures, talking heads, retargeting, and digital characters.

Digital avatars Text-to-motion
553 papers
1.8× vs. 2023
Authors · observed
2,132
Investment · scenario
$576M–$2.30B
Sources and methodology

Publication activity

2019
105
2020
103
2021
125
2022
174
2023
301
2024
463
2025
553

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("motion generation" OR "text to motion" OR "talking head generation" OR "gesture generation" OR "motion retargeting" OR "animatable avatar")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

2,132 distinct OpenAlex Author IDs across all 553 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 553 of 553 papers (100.0%). 14.3% of authorship records lack an Author ID; 11.6% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $1.15B.

  • Creative, music, video content: $4.19B × 10.00%
  • Entertainment: $3.66B × 20.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Creates movement and appearance. Pose estimation and human recognition are covered under vision.

Organizations

Selected examples
  • DeepMotion Product

    SayMotion generates 3D character animation from text.

  • Motorica Product

    Generative character animation.

  • Kinetix Product

    Motion and character animation.

More organizations (8)

World models and simulation

Learned dynamics, action-conditioned generation, interactive environments, and simulation.

World models Action-conditioned simulation
510 papers
3.4× vs. 2023
Authors · observed
1,556
Investment · scenario
$1.41B–$5.63B
Sources and methodology

Publication activity

2019
23
2020
49
2021
77
2022
83
2023
150
2024
231
2025
510

Matches include papers applying these methods. Search precision and recall have not been systematically measured.

Title and abstract matches for articles, preprints and reviews; retracted work is excluded. Versions and overlapping areas may be counted more than once.

("world model" OR "action conditioned generation" OR "interactive simulation") AND (("machine learning" OR "deep learning" OR "neural network" OR "artificial intelligence") OR "reinforcement learning")

Retrieved 2026-09-26.

OpenAlex query results

Publishing authors

1,556 distinct OpenAlex Author IDs across all 510 matching papers. This is an observed count within the query result; no publication-rate assumption is used.

Full query result: 510 of 510 papers (100.0%). 17.2% of authorship records lack an Author ID; 10.2% of retrieved works lack an abstract. Retrieved 2026-09-26.

Applied coauthors are included. Missing or incorrectly linked author records affect these figures. They do not measure research jobs or everyone working in a field.

OpenAlex sample query

Private investment

An assumed allocation of broader investment segments. The range is 0.5–2× the base case. Shares are editorial assumptions, not measured deals or research spending.

Base case: $2.81B.

  • AI infrastructure/models/research/governance: $143.2B × 1.00%
  • Robotics: $7.84B × 10.00%
  • Energy management: $4.64B × 5.00%
  • Entertainment: $3.66B × 10.00%

Company investment, not revenue or research spending. Ranges are scenarios, not confidence intervals.

Stanford AI Index 2026 / Quid

Scope and related research

Model-based reinforcement learning uses these models for decision-making.

Organizations

Selected examples
More organizations (7)
  • Decart Models

    Real-time generative video.

  • NVIDIA Research

    Physical AI and model research.

  • Runway Models

    Video and world models.

  • SpAItial Models

    Spatial foundation models.

  • Tencent / Hunyuan Research

    HY-World reconstructs, generates and simulates 3D environments.

  • Waabi Product

    Autonomous trucking and simulation.

  • Wayve Research

    Research on driving world models, 3D perception and reconstruction.