activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment

Jun Xie, Yingjian Zhu, Feng Chen +9

In this paper, we present our solution for the semi-supervised learning track (MER-SEMI) in MER2025. We propose a comprehensive framework, grounded in the principle that "more is b…

cs.CV2025

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong +4

This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality…

cs.CV2025

Team of One: Cracking Complex Video QA with Model Synergy

Jun Xie, Zhaoran Zhao, Xiongjun Guan +5

We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dat…

cs.CV2025

Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles

Jun Xie, Xiongjun Guan, Yingjian Zhu +5

In this paper, we present the runner-up solution for the Ego4D EgoSchema Challenge at CVPR 2025 (Confirmed on May 20, 2025). Inspired by the success of large models, we evaluate an…

cs.CV2025

PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge

Feng Chen, Kanokphan Lertniphonphan, Qiancheng Yan +4

This report introduces our team's (PCIE_EgoPose) solutions for the EgoExo4D Pose and Proficiency Estimation Challenges at CVPR2025. Focused on the intricate task of estimating 21 3…

cs.CV2025

PCIE_Interaction Solution for Ego4D Social Interaction Challenge

Kanokphan Lertniphonphan, Feng Chen, Junda Xu +4

This report presents our team's PCIE_Interaction solution for the Ego4D Social Interaction Challenge at CVPR 2025, addressing both Looking At Me (LAM) and Talking To Me (TTM) tasks…