9 papers
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Xinkui Zhao, Enbo Chen, Yifan Zhang +4
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However,…
ATOM: Instantiating Budget-Controllable Multi-Agent Collaboration via Nucleus-Electron Hierarchy
Xinkui Zhao, Sai Liu, Yifan Zhang +6
Large Language Model (LLM)-based multi-agent systems rely on optimized collaboration topologies to balance performance and communication costs. However, current methods struggle wi…
DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Flux
Xinkui Zhao, Yifan Zhang, Sai Liu +6
Multi-agent systems (MAS) have demonstrated significant effectiveness in addressing complex problems through coordinated collaboration among heterogeneous agents. However, real-wor…
ProMAS: Proactive Error Forecasting for Multi-Agent Systems Using Markov Transition Dynamics
Xinkui Zhao, Sai Liu, Yifan Zhang +4
The integration of Large Language Models into Multi-Agent Systems (MAS) has enabled the so-lution of complex, long-horizon tasks through collaborative reasoning. However, this coll…
RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
Meng Xi, Sihan Lv, Yechen Jin +4
Retrieval-Augmented Generation (RAG) systems based on Large Language Models (LLMs) have become a core technology for tasks such as question-answering (QA) and content generation. R…
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
Xinkui Zhao, Zuxin Wang, Yifan Zhang +6
The rapid development of multimodal large-language models (MLLMs) has significantly expanded the scope of visual language reasoning, enabling unified systems to interpret and descr…