activity
20242026
most citedVAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

8 citations · 12 across the 15 of their papers we have counts for

collaborators
Showing 2026 · cs.CVShow all

5 papers · 2 filters

cs.CV2026

Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering

Jia Li, Li Dai, Peng Jia +4

In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and…

cs.CV2026

Visual Token Compression Enhances Robustness of MLLMs

Shishen Gu, Jiequan Cui, Wenbo Hu +3

In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…

cs.CV2026

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment

Jia Li, Qian Chen, Wei Wang +5

Personality assessment aims to infer stable traits from dynamic behaviors across modalities like language, voice, and facial expressions. Existing approaches often adopt a uniform…

cs.CV2026

SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

Peng Jia, Zhen Xiao, Jia Li +3

High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-s…

cs.CV2026

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

Jia Li, Yu Zhang, Yin Chen +5

Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations an…