collaborators

14 papers

cs.NI2026

WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks

Zijian Lu, Yiping Zuo, Hao Xu +4

Large language model (LLM) agents are emerging as planners for autonomous wireless network operations. Yet a task answer that is correct at proposal time can still be unsafe at exe…

cs.SE2026

VITAL-RAG: Invariance Race for Context Allocation in Coding Agents

Zijian Lu, Yonghua Lu, Mingcai Chen +4

The paper introduces VITAL-RAG, a method for coding agents that groups retrieved code fragments by their original code object and selectively includes only those that add new task-…

cs.CV2026

Prior Directions: Why GUI Grounding Gets Locked in the Past

Weile Gong, Zijian Lu, Mingcai Chen +3

The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…

cs.RO2026

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…

cs.CV2026

A Comprehensive Survey on World Models for Embodied AI

Xinqing Li, Xin He, Le Zhang +3

Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics,…

cs.CV2026

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

Chenyu Mu, Xin He, Qu Yang +13

Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-f…