From the 1 of 5 linked papers with an AI index.
5 papers
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
Yuanjie Lu, Beichen Wang, Zhengqi Wu +4
The paper introduces APPLV, a system that uses vision‑language models to predict parameters for classical motion planners, combining safety of traditional planners with adaptabilit…
RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation
Zixuan Zhang, Yuqi Chen, Junjie Gao +4
Image-goal navigation is a key challenge in embodied robotics, where an agent must reach a target specified solely by a goal image. While existing reinforcement learning approaches…
Moving Through Clutter: Scaling Data Collection and Benchmarking for 3D Scene-Aware Humanoid Locomotion via Virtual Reality
Beichen Wang, Yuanjie Lu, Linji Wang +2
Recent advances in humanoid locomotion have enabled dynamic behaviors such as dancing, martial arts, and parkour, yet these capabilities are predominantly demonstrated in open, fla…
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
Beichen Wang, Juexiao Zhang, Shuwen Dong +2
Vision Language Models (VLMs) have recently been adopted in robotics for their capability in common sense reasoning and generalizability. Existing work has applied VLMs to generate…
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
Ruixuan Zhang, Beichen Wang, Juexiao Zhang +3
The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which wi…