From the 1 of 15 linked papers with an AI index.
15 papers
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
Xichen Zhang, Guankai Li, Yinghao Zhu +6
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, curren…
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
Shuai Shao, Kangning Zhang, Qingyao Li +7
Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weigh…
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…
OmniGAIA: Towards Native Omni-Modal AI Agents
Xiaoxi Li, Wenxiang Jiao, Jiarui Jin +10
Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However,…
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
Jiarui Jin, Zexuan Yan, Shijian Wang +2
In this paper, we present AgentDisCo, a novel Disentangled and Collaborative agentic architecture that formulates deep research as an adversarial optimization problem between infor…
MMSkills: Towards Multimodal Skills for General Visual Agents
Kangning Zhang, Shuai Shao, Qingyao Li +8
Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable co…