collaborators

10 papers

cs.CV2025

Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation

Yunbo Xu, Xuesong Zhang, Jia Li +2

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has pr…

cs.MM2025

CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition

Yin Chen, Jia Li, Jinpeng Hu +2

Audiovisual emotion recognition (AVER) in the wild is still hindered by pose variation, occlusion, and background noise. Prevailing methods primarily rely on large-scale domain-spe…

cs.CV2025

Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention

Yangche Yu, Yin Chen, Jia Li +6

Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and c…

cs.CL2025

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment

Jia Li, Yang Wang, Wenhao Qian +4

Interview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and co…

cs.CL2025

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors

Jia Li, Yichao He, Jiacheng Xu +4

Accurate and reliable personality assessment plays a vital role in many fields, such as emotional intelligence, mental health diagnostics, and personalized education. Unlike fleeti…

cs.CV2025

Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval

Jian Xiao, Zijie Song, Jialong Hu +4

Recent progress in text-video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes ancho…