6 papers
A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition
Hong Hai Nguyen, Sy Phan Van, Soo-Hyung Kim +1
Facial affect in the wild is naturally multi-task: valence-arousal, discrete expressions, and facial action units describe the same face. Yet real corpora annotate these tasks only…
Faithful Action-unit Causal Reasoning for Counterfactually Faithful Emotion Explanations
Van Thong Huynh, Hong Hai Nguyen, Thuy Pham +2
Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a m…
The Circumplex Degeneracy Behind the Rare-Class Limit in Affect Recognition
Van Thong Huynh, Hong Hai Nguyen, Soo-Hyung Kim
In-the-wild expression recognition persistently fails on a few rare emotions, and the standard explanation is class imbalance. Through a controlled multi-task study on two benchmar…
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
Hoang-Son Vo, Quang-Vinh Nguyen, Seungwon Kim +3
Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synch…
Anatomical Attention Alignment representation for Radiology Report Generation
Quang Vinh Nguyen, Minh Duc Nguyen, Thanh Hoang Son Vo +2
Automated Radiology report generation (RRG) aims at producing detailed descriptions of medical images, reducing radiologists' workload and improving access to high-quality diagnost…
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
Quang Vinh Nguyen, Vo Hoang Thanh Son, Chau Truong Vinh Hoang +3
Naturalistic driving action localization task aims to recognize and comprehend human behaviors and actions from video data captured during real-world driving scenarios. Previous st…