From the 2 of 5 linked papers with an AI index.
5 papers
ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding
Shuo Feng, Piji Li
Embodied Referring Expression Grounding is the task of enabling an agent to navigate in real environments and to localize a remote object based on natural language instructions. In…
Self-Evolving Learning for Embodied AI with Criticality Model
Linxuan He, Yuying Tian, Lingxiang Fan +5
The paper introduces a self‑evolving learning approach for embodied AI that uses a state‑wise criticality model to predict failure and prioritize failure‑prone samples during finet…
DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models
Haoyuan Ji, Lingxiang Fan, Shang Su +4
The paper introduces DC-WAM, a framework that shifts visual supervision in robot world-action models toward dynamic, interaction-relevant features using flow matching and attention…
VPN: Visual Prompt Navigation
Shuo Feng, Zihan Wang, Yuchen Li +6
While natural language is commonly used to guide embodied agents, the inherent ambiguity and verbosity of language often hinder the effectiveness of language-guided navigation in c…
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
Runze Xia, Shuo Feng, Renzhi Wang +3
Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic i…