From the 1 of 23 linked papers with an AI index.
23 papers
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
Shuai Wang, Yaxin Feng, Xuekun Jiang +12
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states a…
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation
Boyu Mi, Mengchen Ma, Yifei Yao +10
The paper introduces REAL, a framework that trains embodied agents for open‑world mobile manipulation using sim‑to‑real consistent environments, hierarchical training, and human‑in…
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
Meng Wei, Chenyang Wan, Xiqian Yu +9
Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instruct…
Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP
Meng Wei, Chenyang Wan, Tai Wang +6
Object-oriented embodied navigation tasks require agents to locate specific objects, either defined by category or images, in unseen environments. While recent methods have made pr…
RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation
Shujie Zhang, Jingkun Yi, Weipeng Zhong +6
Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that ar…
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…