most citedOpen-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

1 citations · 1 across the 5 of their papers we have counts for

collaborators

10 papers

cs.RO2026

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Meng Wei, Chenyang Wan, Xiqian Yu +9

Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instruct…

cs.RO20261 cited

Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

Meng Wei, Chenyang Wan, Tai Wang +6

Object-oriented embodied navigation tasks require agents to locate specific objects, either defined by category or images, in unseen environments. While recent methods have made pr…

cs.CV2026

GTAM: Geometry Grounded Track Anything Model

Chenming Zhu, Peizhou Cao, Jingli Lin +5

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current vi…

cs.CV2026

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

Jinkun Hao, Mingda Jia, Ruiyan Wang +7

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for cont…

cs.CV2026

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Chenming Zhu, Jingli Lin, Yilin Long +4

While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-or…

cs.RO2026

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Yi Chen, Yuying Ge, Hui Zhou +3

The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat…