From the 1 of 14 linked papers with an AI index.
13 papers
Joint On-and-Off Policy Learning for Vision-and-Language Navigation
Qingrong He, Lin Zhao, Kevin Zheng +1
The paper presents JOP-VLN, a three-stage training framework that combines off‑policy imitation learning with on‑policy reinforcement learning to improve navigation and error recov…
Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making
Yue Pei, Hongming Zhang, Chao Gao +7
Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especia…
AdaMorph: Unified Motion Retargeting via Embodiment-Aware Adaptive Transformers
Haoyu Zhang, Shibo Jin, Lusong Li +4
Retargeting human motion to heterogeneous robots is a fundamental challenge in robotics, primarily due to the severe kinematic and dynamic discrepancies between varying embodiments…
Visually-Guided Policy Optimization for Multimodal Reasoning
Zengbin Wang, Feng Xiong, Liang Lin +5
Reinforcement learning with verifiable rewards (RLVR) has significantly advanced the reasoning ability of vision-language models (VLMs). However, the inherent text-dominated nature…
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
Yinghui Li, Jiayi Kuang, Peng Xing +11
Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benc…
AR-MAP: Are Autoregressive Large Language Models Implicit Teachers for Diffusion Large Language Models?
Liang Lin, Feng Xiong, Zengbin Wang +5
Diffusion Large Language Models (DLLMs) have emerged as a powerful alternative to autoregressive models, enabling parallel token generation across multiple positions. However, pref…