8 papers · 2 filters
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
Weiqi Li, Zehao Zhang, Liang Lin +1
Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typicall…
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
Zijian Song, Xiaoxin Lin, Tao Pu +3
Recent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover…
In-Situ Tweedie Discrete Diffusion Models
Xiao Li, Jiaqi Zhang, Shuxiang Zhang +3
While diffusion models excel at generating continuous data such as images, adapting them to discrete tasks has relied on indirect approaches that either operate in continuous embed…
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
Zijian Song, Sihan Qin, Tianshui Chen +2
The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation mo…
GS: Generative Segmentation via Label Diffusion
Yuhao Chen, Shubin Chen, Liang Lin +1
Language-driven image segmentation is a fundamental task in vision-language understanding, requiring models to segment regions of an image corresponding to natural language express…
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
Weiqi Li, Zehao Zhang, Liang Lin +1
\textbf{Synthetic human dynamics} aims to generate photorealistic videos of human subjects performing expressive, intention-driven motions. However, current approaches face two cor…