10 papers
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
Yunbo Xu, Xuesong Zhang, Jia Li +2
Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has pr…
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
Yin Chen, Jia Li, Jinpeng Hu +2
Audiovisual emotion recognition (AVER) in the wild is still hindered by pose variation, occlusion, and background noise. Prevailing methods primarily rely on large-scale domain-spe…
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
Yangche Yu, Yin Chen, Jia Li +6
Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and c…
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
Jia Li, Yang Wang, Wenhao Qian +4
Interview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and co…
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
Jia Li, Yichao He, Jiacheng Xu +4
Accurate and reliable personality assessment plays a vital role in many fields, such as emotional intelligence, mental health diagnostics, and personalized education. Unlike fleeti…
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
Jian Xiao, Zijie Song, Jialong Hu +4
Recent progress in text-video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes ancho…