works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

Mingxin Wang, Bin Hu, Bin Qian +12

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future sup…

cs.CV2026

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Ruiyan Gong, Yingnan Guo, Junjun Hu +44

The paper presents ABot-N1, a visual‑language navigation foundation model that separates high‑level reasoning from low‑level control via a slow‑fast architecture and pixel‑based go…

cs.CV2026

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Ronghan Chen, Yandan Yang, Zuojin Tang +18

Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…

cs.CV2026

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

Qiming Li, Tianlun Li, Xiaolong Cheng +5

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, exis…

cs.RO2026

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

Ruiyan Gong, Meisheng Zhang, Yuxiang Zhao +12

Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigat…

cs.CV2026

ABot-OCR Technical Report

Kaitao Jiang, Ruiyan Gong, Xiaolong Cheng +3

We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely…