activity
20242026
collaborators

6 papers

cs.AI2026

To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models

Haoqing Wang, Xiang Long, Ziheng Li +3

Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level…

cs.RO2025

MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis

Xiaowei Chi, Kuangzhi Ge, Jiaming Liu +9

Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, cur…

cs.RO2025

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

Weifan Zhang, Tingguang Li, Yuzhen Liu

Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework…

cs.RO2025

VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation

Kefeng Huang, Tingguang Li, Yuzhen Liu +3

Diffusion policy has demonstrated promising performance in the field of robotic manipulation. However, its effectiveness has been primarily limited in short-horizon tasks, and its…

cs.RO2024

VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation

Bangguo Yu, Yuzhen Liu, Lei Han +3

Following human instructions to explore and search for a specified target in an unfamiliar environment is a crucial skill for mobile service robots. Most of the previous works on o…

cs.CV2024

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Sijie Cheng, Kechen Fang, Yangyang Yu +6

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoTh…