16 papers
Motion Attribution for Video Generation
Xindi Wu, Despoina Paschalidou, Jun Gao +5
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a m…
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
Xiaobin Rong, Jun Gao, Zheng Wang +3
Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but h…
OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics
Zhuoyuan Wu, Jun Gao
We present OSCAR, a precise action-conditioned video world model that generalizes across different robot embodiments and enables robot policy evaluation. Existing video world model…
AFUN: Towards an Affordance Foundation Model for Functionality Understanding
Zhaoning Wang, Yi Zhong, Jiawei Fu +2
Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstructured real-world environments.…
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
Bingyu Li, Da Zhang, Tao Huo +3
Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temporal visual reasoning remains…
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Fangfu Liu, Kai He, Tianchang Shen +7
World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many gen…