14 papers
WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks
Zijian Lu, Yiping Zuo, Hao Xu +4
Large language model (LLM) agents are emerging as planners for autonomous wireless network operations. Yet a task answer that is correct at proposal time can still be unsafe at exe…
VITAL-RAG: Invariance Race for Context Allocation in Coding Agents
Zijian Lu, Yonghua Lu, Mingcai Chen +4
The paper introduces VITAL-RAG, a method for coding agents that groups retrieved code fragments by their original code object and selectively includes only those that add new task-…
Prior Directions: Why GUI Grounding Gets Locked in the Past
Weile Gong, Zijian Lu, Mingcai Chen +3
The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
A Comprehensive Survey on World Models for Embodied AI
Xinqing Li, Xin He, Le Zhang +3
Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics,…
The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation
Chenyu Mu, Xin He, Qu Yang +13
Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-f…