4 papers
AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding
Yuxiang Duan, Huining Li, Ao Li +6
Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting e…
M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities
Ao Li, Jinghui Zhang, Luyu Li +10
As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex…
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
Yuxiang Duan, Ao Li, Yingqin Li +2
Multimodal large language models (MLLMs) have shown remarkable capabilities in a wide range of vision-language tasks. However, the large number of visual tokens introduces signific…