7 papers
GraphIR: Architecture-Level Search States for LLM-Guided Neural Architecture Evolution
Zhen Liu, Wanqi Zhou, Shuanghao Bai +3
Large language models (LLMs) enable neural architecture search (NAS) directly over executable neural network programs. However, code-level flexibility does not provide the architec…
Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search
Zhen Liu, Yuhan Liu, Jinjun Wang +3
This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring new designs under expensive evaluations.…
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Kangyi Wu, Pengna Li, Kailin Lyu +5
Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Large Language Models(Video-LLM…
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
Zhen Liu, Yuhan Liu, Jinjun Wang +3
In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends on balancing behavioral div…
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation
Zhen Liu, Yuhan Liu, Jinjun Wang +3
Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dynamic entanglement between la…
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
Kangyi Wu, Pengna Li, Jingwen Fu +4
Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However…