1 paper
Boyu Chen, Zikang Wang, Zhengrong Yue +9
By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and no…