28 citations · 53 across the 18 of their papers we have counts for
18 papers
Depth-aware Test-Time Training for Zero-shot Video Object Segmentation
Weihuang Liu, Xi Shen, Haolun Li +4
Zero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model…
EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face Generation
Guanwen Feng, Haoran Cheng, Yunan Li +5
Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately a…
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
Lian Huang, Chi-Man Pun
Due to the successful application of deep learning, audio spoofing detection has made significant progress. Spoofed audio with speech synthesis or voice conversion can be well dete…
COMMA: Co-Articulated Multi-Modal Learning
Lianyu Hu, Liqing Gao, Zekang Liu +2
Pretrained large-scale vision-language models such as CLIP have demonstrated excellent generalizability over a series of downstream tasks. However, they are sensitive to the variat…
ELF: An End-to-end Local and Global Multimodal Fusion Framework for Glaucoma Grading
Wenyun Li, Chi-Man Pun
Glaucoma is a chronic neurodegenerative condition that can lead to blindness. Early detection and curing are very important in stopping the disease from getting worse for glaucoma…
Perceptual MAE for Image Manipulation Localization: A High-level Vision Learner Focusing on Low-level Features
Xiaochen Ma, Jizhe Zhou, Xiong Xu +2
Nowadays, multimedia forensics faces unprecedented challenges due to the rapid advancement of multimedia generation technology thereby making Image Manipulation Localization (IML)…