activity
20242026
most citedVAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

8 citations · 12 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Visual Token Compression Enhances Robustness of MLLMs

Shishen Gu, Jiequan Cui, Wenbo Hu +3

In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…

cs.CV2026

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment

Jia Li, Qian Chen, Wei Wang +5

Personality assessment aims to infer stable traits from dynamic behaviors across modalities like language, voice, and facial expressions. Existing approaches often adopt a uniform…

cs.CV2026

SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

Peng Jia, Zhen Xiao, Jia Li +3

High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-s…

cs.CV2026

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

Jia Li, Yu Zhang, Yin Chen +5

Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations an…

cs.CV2025

Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation

Yunbo Xu, Xuesong Zhang, Jia Li +2

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has pr…

cs.CV2025

Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention

Yangche Yu, Yin Chen, Jia Li +6

Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and c…