works on

From the 1 of 24 linked papers with an AI index.

activity
20242026
most citedOpen-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

Boyu Mi, Mengchen Ma, Yifei Yao +10

The paper introduces REAL, a framework that trains embodied agents for open‑world mobile manipulation using sim‑to‑real consistent environments, hierarchical training, and human‑in…

cs.CV2026

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

Jiaxin Liu, Xun Xu, Zhenhao Zhang +5

Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-…

cs.CV2026

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

Weipeng Zhong, Peizhou Cao, Yichen Jin +9

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typic…

cs.CV2025

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

Sihao Lin, Zerui Li, Xunyi Zhao +10

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…

cs.CV2025

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang +2

3D visual grounding (3DVG) is challenging due to the need to understand 3D spatial relations. While supervised approaches have achieved superior performance, they are constrained b…