10 papers
HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition
Kunjie Deng, Jinghui Zhang, Weidong Chen +4
Surgical video phase recognition is a fundamental task in computer-assisted intervention, supporting workflow understanding, intraoperative guidance, and surgical quality assessmen…
Geometry-aware Gaussian Prior and Axial Attention for Cervical Cytology Image Classification
Yating Li, Cheng Ye, Nenan Lyu +2
Accurate cervical cytology image classification is a key component of automated cervical cancer screening, where reliable recognition of normal, precancerous, and cancer-associated…
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
Dexiang Hong, Yijie Guo, Weidong Chen +4
Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main…
Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning
Zihan Meng, Dexiang Hong, Weidong Chen +3
Audio-visual captioning generates natural language descriptions from video and audio content. Multimodal LLMs have advanced this task, but both modalities contribute many tokens to…
Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction
Weidong Chen, Cheng Ye, Zhendong Mao +3
Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic…
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
Liping Wang, Cheng Ye, Weidong Chen +3
Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Existing approaches usually rel…