13 papers
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Mingxin Wang, Bin Hu, Bin Qian +12
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future sup…
BasketEvent: Understanding Who Did What and When in Basketball Videos
Yu Zhang, Jiayuan Rao, Haoning Wu +1
Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears. However, exist- ing metho…
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Ronghan Chen, Yandan Yang, Zuojin Tang +18
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…
PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
Bin Hu, Yanwen Ma, Jiehui Huang +14
Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…
Count Anything at Any Granularity
Chang Liu, Haoning Wu, Weidi Xie
Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that…
Improving Human Image Animation via Semantic Representation Alignment
Chang Liu, Mengting Chen, Yixuan Huang +5
The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long…