activity
20242026
most citedRobobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

1 citations · 1 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment

Huaihai Lyu, Chaofan Chen, Yuheng Ji +4

We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action representations compatible with the s…

cs.CV2026

AoE: Always-on Egocentric Human Video Collection for Embodied AI

Bowen Yang, Zishuo Li, Yang Sun +15

Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high in…

cs.CV2025

OmniSAT: Compact Action Token, Faster Auto Regression

Huaihai Lyu, Chaofan Chen, Senwei Xie +4

Existing Vision-Language-Action (VLA) models can be broadly categorized into diffusion-based and auto-regressive (AR) approaches: diffusion models capture continuous action distrib…

cs.CV2025

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

Huajie Tan, Yuheng Ji, Xiaoshuai Hao +4

Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI). Exis…

cs.CV2025

MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles

Yuheng Ji, Huajie Tan, Cheng Chi +8

We introduce \textsc{MathSticks}, a benchmark for Visual Symbolic Compositional Reasoning (VSCR), which unifies visual perception, symbolic manipulation, and arithmetic consistency…

cs.CV2025

SafeMap: Robust HD Map Construction from Incomplete Observations

Xiaoshuai Hao, Lingdong Kong, Rong Yin +4

Robust high-definition (HD) map construction is vital for autonomous driving, yet existing methods often struggle with incomplete multi-view camera data. This paper presents SafeMa…