13 papers
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
Zhipeng Bao, Zhen Zhu, Nupur Kumari +4
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study…
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
Junjie Ye, Rong Xue, Basile Van Hoorick +4
The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limi…
Robot Critics that Sweat the Small Stuff
Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan +3
Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. Ho…
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
Junjie Ye, Rong Xue, Basile Van Hoorick +6
Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While vide…
Capturing Visual Environment Structure Correlates with Control Performance
Jiahua Dong, Yunze Man, Pavel Tokmakov +1
The choice of visual representation is key to scaling generalist robot policies. However, direct evaluation via policy rollouts is expensive, even in simulation. Existing proxy met…
AnyView: Synthesizing Any Novel View in Dynamic Scenes
Basile Van Hoorick, Dian Chen, Shun Iwase +7
Modern generative video models excel at producing convincing, high-quality outputs, but struggle to maintain multi-view and spatiotemporal consistency in highly dynamic real-world…