8 citations · 12 across the 15 of their papers we have counts for
5 papers · 2 filters
Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering
Jia Li, Li Dai, Peng Jia +4
In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and…
Visual Token Compression Enhances Robustness of MLLMs
Shishen Gu, Jiequan Cui, Wenbo Hu +3
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…
Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment
Jia Li, Qian Chen, Wei Wang +5
Personality assessment aims to infer stable traits from dynamic behaviors across modalities like language, voice, and facial expressions. Existing approaches often adopt a uniform…
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
Peng Jia, Zhen Xiao, Jia Li +3
High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-s…
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
Jia Li, Yu Zhang, Yin Chen +5
Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations an…