From the 1 of 115 linked papers with an AI index.
115 papers
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Jingtao Sun, Xiaohai He, Yike Zhang +4
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a…
Sensorimotor Stickies: A Reconfigurable On-Body Platform for Closed-Loop Sensorimotor Training
Tianhong Catherine Yu, Jiwei Zheng, Chi-Jung Lee +7
Closed-loop sensorimotor training systems can improve learning by sensing movement and delivering real-time feedback, yet most are built as fixed implementations tied to a single t…
MANGO-Grasp: Mahalanobis Fields over Geometry-Oriented 3D Gaussians for Cross-Embodiment Dexterous Grasping
Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou +2
Cross-embodiment dexterous grasping aims to synthesize stable grasps across heterogeneous multi-fingered hands with little or no embodiment-specific tuning. Existing interaction-ce…
Hermite Curves as Trajectory Priors for Vision-Language-Action Models
Qi Lv, Jianming Xing, Zhao Yang +5
Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten eac…
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
Neel Mokaria, Rishie Raj, Dheeraj Baiju +13
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Quanjian Song, Yiren Song, Kelly Peng +2
WorldWander is a framework that translates video content between first‑person (egocentric) and third‑person (exocentric) views using video diffusion transformers and in‑context lea…