From the 1 of 7 linked papers with an AI index.
7 papers
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Mingxin Wang, Bin Hu, Bin Qian +12
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future sup…
ABot-N1: Toward a General Visual Language Navigation Foundation Model
Ruiyan Gong, Yingnan Guo, Junjun Hu +44
The paper presents ABot-N1, a visual‑language navigation foundation model that separates high‑level reasoning from low‑level control via a slow‑fast architecture and pixel‑based go…
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Ronghan Chen, Yandan Yang, Zuojin Tang +18
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
Qiming Li, Tianlun Li, Xiaolong Cheng +5
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, exis…
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
Ruiyan Gong, Meisheng Zhang, Yuxiang Zhao +12
Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigat…
ABot-OCR Technical Report
Kaitao Jiang, Ruiyan Gong, Xiaolong Cheng +3
We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely…