1 citations · 1 across the 16 of their papers we have counts for
8 papers · 1 filter
LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment
Huaihai Lyu, Chaofan Chen, Yuheng Ji +4
We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action representations compatible with the s…
AoE: Always-on Egocentric Human Video Collection for Embodied AI
Bowen Yang, Zishuo Li, Yang Sun +15
Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high in…
OmniSAT: Compact Action Token, Faster Auto Regression
Huaihai Lyu, Chaofan Chen, Senwei Xie +4
Existing Vision-Language-Action (VLA) models can be broadly categorized into diffusion-based and auto-regressive (AR) approaches: diffusion models capture continuous action distrib…
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
Huajie Tan, Yuheng Ji, Xiaoshuai Hao +4
Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI). Exis…
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
Yuheng Ji, Huajie Tan, Cheng Chi +8
We introduce \textsc{MathSticks}, a benchmark for Visual Symbolic Compositional Reasoning (VSCR), which unifies visual perception, symbolic manipulation, and arithmetic consistency…
SafeMap: Robust HD Map Construction from Incomplete Observations
Xiaoshuai Hao, Lingdong Kong, Rong Yin +4
Robust high-definition (HD) map construction is vital for autonomous driving, yet existing methods often struggle with incomplete multi-view camera data. This paper presents SafeMa…