10 papers
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
Yuanzhi Xu, Qian Gao, Jun Fan +4
The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to…
RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation
Sixu Lin, Junliang Chen, Huaiyuan Xu +8
Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predictive flow planners to guide 3D…
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
Sixu Lin, Yunpeng Qing, Litao Liu +4
Recent progress in Reinforcement Learning (RL) provides a principled approach to optimizing Vision-Language-Action (VLA) models, facilitating a shift from trajectory imitation to a…
BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning
Yunpeng Qing, Yixiao Chi, Shuo Chen +5
Recent advances in offline Reinforcement Learning (RL) have proven that effective policy learning can benefit from imposing conservative constraints on pre-collected datasets. Howe…
SignBot: Learning Human-to-Humanoid Sign Language Interaction
Guanren Qiao, Sixu Lin, Ronglai Zuo +3
Sign language is a natural and visual form of language that uses movements and expressions to convey meaning, serving as a crucial means of communication for individuals who are de…