From the 1 of 23 linked papers with an AI index.
23 papers
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Haodong Li, Shaoteng Liu, Tianyu Wang +7
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time…
Reflecting Process Expertise in Procedural Material Generation
Kunal Gupta, Gaurav Joshi, Yen-Ru Chen +3
The paper introduces a method that captures expert material‑creation workflows as textual process traces and uses large language models to generate and compile these traces into ed…
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
Hanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu +2
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional.…
-Scene: Physically Grounded Image-to-3D Scene Reconstruction
Haodong Li, Lulu Shao, Haolin Lu +4
Reconstructing compositional 3D scenes from a single image is a fundamental challenge in 3D world modeling. Recent methods can recover high-fidelity, complete 3D objects and predic…
Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos
Anurag Ghosh, Francesco Pittaluga, Khiem Vuong +4
Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations…
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
Abhishek Aich, Sparsh Garg, Vijay Kumar BG +2
Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet verification remains sparse: man…