3 papers
cs.CV2026
Entropy Reveals Block Importance in Masked Self-Supervised Vision Transformers
Peihao Xiang, Kaida Wu, Ou Bai
Masked self-supervised vision transformers have become a dominant pretraining paradigm, yet their substantial model size poses significant challenges for resource-constrained deplo…
cs.SD2025
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
Chaohao Lin, Xu Zheng, Kaida Wu +2
Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are…
cs.CV2024
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition
Peihao Xiang, Kaida Wu, Ou Bai
This paper expands the cascaded network branch of the autoencoder-based multi-task learning (MTL) framework for dynamic facial expression recognition, namely Multi-Task Cascaded Au…