10 papers
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
Liangyu Fu, Junbo Wang, Yuke Li +3
Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap bet…
Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning
Junbo Wang, Liangyu Fu, Yuke Li +2
Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive subtle, ambiguous, and tempora…
EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
Junbo Wang, Liangyu Fu, Yuke Li +3
Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Transformer-based methods, have a…
DiffVC: A Non-autoregressive Framework Based on Diffusion Model for Video Captioning
Junbo Wang, Liangyu Fu, Yuke Li +4
Current video captioning methods usually use an encoder-decoder structure to generate text autoregressively. However, autoregressive methods have inherent limitations such as slow…
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
Fengjian Xue, Xuecheng Wu, Heli Sun +8
Facial expression image editing requires fine-grained control to strictly preserve human identity and background while precisely manipulating expression. However, existing editing…
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
Xuecheng Wu, Junxiao Xue, Xinyi Yin +7
Affective video facial analysis (AVFA) has emerged as a key research field for building emotion-aware intelligent systems, yet this field continues to suffer from limited data avai…