Workshop on 4D World Models: Bridging Generation and Reconstruction @ CVPR 2026
Workshop on 4D World Models: Bridging Generation and Reconstruction @ CVPR 2026 —
129 open resources, workshops, research organizations, and technical reports from the curated list.
129 resources
Try a model name, organization, or arXiv ID.
Workshop on 4D World Models: Bridging Generation and Reconstruction @ CVPR 2026 —
2nd Workshop on World Models @ ICLR 2026 —
Workshop on World Modeling @ Mila 2026 —
1st Workshop on Benchmarking World Models.
OpenDriveLab World Model Track @ CVPR 2025 —
OpenDriveLab Predictive World Model Track @ CVPR 2024 —
Argoverse 3D Occupancy Forecasting @ CVPR 2023 —
Sydney; latent dynamics, generative simulation, evaluation, planning/control; co-located AV Causal Reasoning Retrieval Challenge.
world models and WAMs for robot reasoning, learning, and evaluation.
Sydney; world models that keep learning after deployment from observation, memory, feedback, and interaction.
Atlanta; patient world models, intervention-aware reasoning, and clinical trial simulation as a falsifiable world-model testbed.
Broad cross-domain curation
Awesome World Models (knightnemo)
General video generation, embodied AI, AD
Awesome World Models (leofan90)
Driving-specific papers, benchmarks, challenges
Robotics, embodied AI, VLA-adjacent work
Curated trajectory from video gen to world modeling
Autoregressive diffusion recipes for scalable, consistent, interactive video world models
Interactive video world modeling papers, benchmarks, datasets, and resources
Physical AI: VLA models, world models, embodied robotic foundations
World Action Models: survey, taxonomy, papers, data, and evaluation resources
Survey companion: understanding world or predicting future?
Physics plausibility in video world models
Robustness-focused driving world models
Companion repo for From Masks to Worlds; emphasizes evolutionary roadmaps and memory-augmented world models
Survey-centered repo for a broad AI view of world models
Comprehensive embodied AI + world model papers
Physical simulation + world models for embodied AI
Provides remote, validated physical experiments and recorded action-state feedback for wireless-network world-model development; an experimentation platform rather than a completed autonomous world model.
Inspectable video-data recipes with provenance and controlled pretraining studies on Wan 2.1 and V-JEPA 2.1; adjacent infrastructure for predictive pretraining.
Unitree framework that predicts future robot interactions for visual simulation and policy enhancement across embodiments.
World foundation model platform for Physical AI (robots + AD); open-weight under permissive license
Open omnimodal WFM unifying reasoning, world generation, simulation, and action modeling
High-performance inference and serving library for interactive autoregressive video and world models
Next-gen Cosmos WFM: flow-based, unifies Text/Image/Video2World; open checkpoints
Reproducible world-model research platform with data layer, baselines, planners, and OOD tasks
Full-stack framework for building real-time interactive video world models from open video backbones
Minimalist future-video-prediction codebase with configs, eval scripts, and checkpoints
Open minute-scale, 720p video world model with 6-DoF camera control
Open recipe for real-time autoregressive video diffusion world rollouts
Open real-time long-video generation stack relevant to live world modeling
Open-source toolkit for driving world models (SenseTime)
Open interactive game world model stack (SkyworkAI)
Open robotic manipulation world foundation platform
Large-scale manipulation platform and dataset for embodied world models
Gaussian world model codebase for robotic manipulation
Open 3D world generation / simulation stack (Tencent)
Text/image-to-3D explorable world generation (Tencent)
Reference implementation of the Dreamer family
Open-source TD-MPC2 codebase, 104 tasks
Meta's latest JEPA world model for video understanding and robotic planning
Meta's original JEPA implementations
Open reproduction of Oasis Minecraft world model
Open-source WAM stack with checkpoints, eval tooling, and embodiment adaptation scripts
AgiBot's action-conditional embodied world model
Scalable multi-agent multi-view video world model
WorldLens benchmark dataset + leaderboard
Diffusion-based Atari world model + RL agent
Open-source general world simulator with real-time interactivity
AMD open-source interactive world model for game-like environments
Lean, provable JEPA self-supervised training framework (SIGReg); ~50-line core
Full-spectrum driving world model evaluation
Official embodied world model leaderboard
Central HuggingFace hub for videogen, occgen, and lidargen resources
Physics-centric benchmark dataset for world models and VLMs
Benchmark dataset for judging video generation models as world models
Multi-turn interactive video world model evaluation
Long-horizon stability for interactive world models (action, vision, physics, memory). Paper: WorldOdysseyBench in Benchmarks.
Multi-view spatial-consistency benchmark for world models
Official Meta V-JEPA 2 checkpoints (ViT-L/H/G)
Open WebWorld-8B/14B/32B web world-model checkpoints for agent training and lookahead search
Unreal Engine pipeline separates physics trajectory collection from offline rendering to produce action-aligned multi-view video, with distributed scene screening and recovery; reports 8,767 hours across 1080p and 720p outputs.
Gameplay UI taxonomy, paired synthetic videos, and in-the-wild evaluation clips support temporally consistent HUD removal; a controlled pilot evaluates cleaned footage for world-model training.
SoundSpaces-based open platform generates continuous binaural audio along simulated navigation rollouts, supplying reproducible acoustic training data for audio-based world models.
Official dataset behind WorldArena embodied evaluation
Replay-grounded egocentric Counter-Strike trajectories with video, actions, states, events, and language
UE5 replay dataset for physics-editable world models with gravity interventions, actions, states, and multimodal rollouts
Large-scale semantic world-model dataset for mobile GUI agents
Highly dynamic UAV-view dataset for world models
Multi-domain, multi-modal 4D world modeling dataset and benchmark
Large-scale egocentric human dataset for robot learning and transfer
Real-world interactive world-model data used by MagicWorld-style exploration
Spatial-consistency benchmark data for memory-aided world models
Microscale simulation data and rubric-based benchmark
Public RoboTwin/LIBERO evaluation packages for action-conditioned world-model faithfulness probing
Who is building what, restricted to public, primary-source material already linked in this list. "Representative entries" point to sections where the full citations and badges live. Claims about unpublished internal systems are deliberately excluded.
World foundation model platform for Physical AI: video WFMs, driving data engines, real-time closed-loop simulation, omnimodal Cosmos 3
Foundation world models for playable environments; generalist 3D agents
Non-generative predictive representation program; video JEPA world models with robot planning
Generative driving world models with fine-grained controllability
Spatially grounded multimodal 3D world generation
Humanoid robotics; world models for real-robot video prediction and policy evaluation
Robotic manipulation world platforms, embodied data engines, embodied evaluation
Open real-time interactive game/world stacks; 3D world generation
Explorable, mesh-based, and simulatable 3D world generation
Video generation framed as world simulation — the claim that started the 2024 debate
Occupancy networks for planning, presented publicly in engineering talks (CVPR 2022 WAD keynote; no citable primary paper)
Representative entries: Xiaomi EV World Model, Xiaomi-Robotics-U0, MiLA, DGGT.
Representative entries: MineWorld, Latent Spatial Memory.
Representative entries: WorldVLA, Qwen-RobotWorld, Qwen-AgentWorld, WorldOlympiad.
Representative entries: OpenDWM, MaskGWM, UniMLVG.
Representative entries: ReconDreamer, GigaWorld, GigaBrain.
Representative entries: WBench.
Representative entries: LingBot-VA, LingBot-World.
This section is intentionally selective. It favors official research-lab posts, technical reports, and a small number of high-signal explainers over general-audience trend pieces.
Research preview of action- and camera-controlled audio-video worlds; persistent context and timestamped interactions.
Spatial generation, 3D outputs, and space-time simulation; select-partner early access.
(original explainer site)
World Models (original explainer site)