10 papers
Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
Seongro Yoon, Donghyeon Cho, Jinsun Park +1
Understanding facial expressions in videos requires modeling subtle and localized facial dynamics under unconstrained conditions. Although recent Vision Transformer (ViT)-based vid…
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
Nishit Poddar, Aglind Reka, Diana-Laura Borza +4
Micro-actions, fleeting and low-amplitude motions, such as glances, nods, or minor posture shifts, carry rich social meaning but remain difficult for current action recognition mod…
LIA-X: Interpretable Latent Portrait Animator
Yaohui Wang, Di Yang, Xinyuan Chen +3
We introduce LIA-X, a novel interpretable portrait animator designed to transfer facial dynamics from a driving video to a source portrait with fine-grained control. LIA-X is an au…
Just Dance with ! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
Snehashis Majhi, Giacomo D'Amicantonio, Antitza Dantcheva +5
Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-worl…
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
Junuk Cha, Seongro Yoon, Valeriya Strizhkova +2
3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is…
CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets
Tanay Agrawal, Mohammed Guermal, Michal Balazia +1
Challenges in cross-learning involve inhomogeneous or even inadequate amount of training data and lack of resources for retraining large pretrained models. Inspired by transfer lea…