5 papers
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Mingxin Wang, Bin Hu, Bin Qian +12
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future sup…
MobileSAM2: Lightweight Segment Anything for Spatial Intelligence
Kai Jiang, Jiaxing Huang, Jingyi Zhang +5
The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such…
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
Qiming Li, Tianlun Li, Xiaolong Cheng +5
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, exis…
A Survey on Vision Autoregressive Model
Kai Jiang, Jiaxing Huang
Autoregressive models have demonstrated great performance in natural language processing (NLP) with impressive scalability, adaptability and generalizability. Inspired by their not…
Open-Vocabulary Object Detection via Language Hierarchy
Jiaxing Huang, Jingyi Zhang, Kai Jiang +1
Recent studies on generalizable object detection have attracted increasing attention with additional weak supervision from large-scale datasets with image-level labels. However, we…