most citedRoboDreamer: Learning Compositional World Models for Robot Imagination

4 citations · 6 across the 2 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation

Jiaben Chen, Sixun Dong, Qinhong Zhou +4

Long-form narrative-to-film generation requires shot-level controllability and cross-clip consistency in both visual identity and character behavior-requirements that remain diffic…

cs.CV2024

Compositional Physical Reasoning of Objects and Events from Videos

Zhenfang Chen, Shilong Dong, Kexin Yi +5

Understanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and sha…

cs.CV2024

SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge

Andong Wang, Bo Wu, Sunli Chen +5

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks…

cs.LG2024

LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery

Pingchuan Ma, Tsun-Hsuan Wang, Minghao Guo +5

Large Language Models have recently gained significant attention in scientific discovery for their extensive knowledge and advanced reasoning capabilities. However, they encounter…

cs.CV2024

Physically Compatible 3D Object Modeling from a Single Image

Minghao Guo, Bohan Wang, Pingchuan Ma +6

We present a computational framework that transforms single images into 3D physical objects. The visual geometry of a physical object in an image is determined by three orthogonal…

cs.RO20244 cited

RoboDreamer: Learning Compositional World Models for Robot Imagination

Siyuan Zhou, Yilun Du, Jiaben Chen +3

Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environme…