8 citations · 12 across the 15 of their papers we have counts for
11 papers · 1 filter
Visual Token Compression Enhances Robustness of MLLMs
Shishen Gu, Jiequan Cui, Wenbo Hu +3
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…
Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment
Jia Li, Qian Chen, Wei Wang +5
Personality assessment aims to infer stable traits from dynamic behaviors across modalities like language, voice, and facial expressions. Existing approaches often adopt a uniform…
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
Peng Jia, Zhen Xiao, Jia Li +3
High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-s…
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
Jia Li, Yu Zhang, Yin Chen +5
Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations an…
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
Yunbo Xu, Xuesong Zhang, Jia Li +2
Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has pr…
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
Yangche Yu, Yin Chen, Jia Li +6
Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and c…