activity
20232026
most citedHiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition

70 citations · 133 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Hui Wei, Licai Sun, Guoying Zhao

Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art…

cs.CV2025

Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

Licai Sun, Xingxun Jiang, Haoyu Chen +7

Current facial emotion recognition systems are predominately trained to predict a fixed set of predefined categories or abstract dimensional values. This constrained form of superv…

cs.CV2025

MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric Guidance

Mengting Wei, Yante Li, Tuomas Varanka +2

In this study, we propose a method for video face reenactment that integrates a 3D face parametric model into a latent diffusion framework, aiming to improve shape consistency and…

cs.CV2024

Multimodal Fusion with Pre-Trained Model Features in Affective Behaviour Analysis In-the-wild

Zhuofan Wen, Fengyu Zhang, Siyuan Zhang +6

Multimodal fusion is a significant method for most multimodal tasks. With the recent surge in the number of large pre-trained models, combining both multimodal fusion methods and p…

cs.CV2024★ 70 cited

HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition

Licai Sun, Zheng Lian, Bin Liu +1

Audio-Visual Emotion Recognition (AVER) has garnered increasing attention in recent years for its critical role in creating emotion-ware intelligent machines. Previous efforts in t…

cs.CV2024★ 29 cited

SVFAP: Self-supervised Video Facial Affect Perceiver

Licai Sun, Zheng Lian, Kexin Wang +5

Video-based facial affect analysis has recently attracted increasing attention owing to its critical role in human-computer interaction. Previous studies mainly focus on developing…