3 papers
cs.CV2026
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
Kaiwen Xue, Tao Wei, Guoxin Zhang +5
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluat…
cs.LG2026
Skill Reuse as Compression in Agentic RL
Zhikun Xu, Yu Feng, Jacob Dineen +3
Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successfu…
cs.RO2026
FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching
Zihao Zheng, Xingyue Zhou, Zhihao Mao +7
Vision-Language-Navigation (VLN) models exhibit excellent navigation accuracy but incur high computational overhead. Token caching has emerged as a promising training-free strategy…