5 papers · 1 filter
SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks
Tao Yu, Yifei Qu, Zhiqing Cui +14
Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA e…
More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment
Jun Xie, Yingjian Zhu, Feng Chen +9
In this paper, we present our solution for the semi-supervised learning track (MER-SEMI) in MER2025. We propose a comprehensive framework, grounded in the principle that "more is b…
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
Zhepeng Wang, Yingjian Zhu, Guanghao Dong +4
This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality…
Team of One: Cracking Complex Video QA with Model Synergy
Jun Xie, Zhaoran Zhao, Xiongjun Guan +5
We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dat…
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles
Jun Xie, Xiongjun Guan, Yingjian Zhu +5
In this paper, we present the runner-up solution for the Ego4D EgoSchema Challenge at CVPR 2025 (Confirmed on May 20, 2025). Inspired by the success of large models, we evaluate an…