From the 1 of 23 linked papers with an AI index.
1 citations · 1 across the 14 of their papers we have counts for
24 papers · 1 filter
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Haodong Li, Shaoteng Liu, Tianyu Wang +7
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time…
Reflecting Process Expertise in Procedural Material Generation
Kunal Gupta, Gaurav Joshi, Yen-Ru Chen +3
The paper introduces a method that captures expert material‑creation workflows as textual process traces and uses large language models to generate and compile these traces into ed…
-Scene: Physically Grounded Image-to-3D Scene Reconstruction
Haodong Li, Lulu Shao, Haolin Lu +4
Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that limit physical validity and down…
Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos
Anurag Ghosh, Francesco Pittaluga, Khiem Vuong +4
Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations…
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
Abhishek Aich, Sparsh Garg, Vijay Kumar BG +2
Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet verification remains sparse: man…
Materialist: Physically Based Editing Using Single-Image Inverse Rendering
Lezhong Wang, Duc Minh Tran, Ruiqi Cui +5
Achieving physically consistent image editing remains a significant challenge in computer vision. Existing image editing methods typically rely on neural networks, which struggle t…