From the 1 of 13 linked papers with an AI index.
9 papers · 1 filter
Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation
Xunyi Zhao, Sihao Lin, Gengze Zhou +5
Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an under-specified natural-language description. Such ambiguity…
IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
Lin Qian, Shijie Li, Sihao Lin +4
Existing object navigation benchmarks usually tell an embodied agent which object category to find, such as microwave or chair. Human-facing embodied AI is often asked something le…
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
Sihao Lin, Zerui Li, Xunyi Zhao +10
Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
Changlin Li, Jiawei Zhang, Shuhao Liu +4
Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training th…
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba
Xiuwei Chen, Wentao Hu, Xiao Dong +7
Transformer-based architectures have become the backbone of both uni-modal and multi-modal foundation models, largely due to their scalability via attention mechanisms, resulting i…
Learning A Zero-shot Occupancy Network from Vision Foundation Models via Self-supervised Adaptation
Sihao Lin, Daqi Liu, Ruochong Fu +6
Estimating the 3D world from 2D monocular images is a fundamental yet challenging task due to the labour-intensive nature of 3D annotations. To simplify label acquisition, this wor…