6 papers
Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning
Zhihua Xu, Zhijing Yang, Yufeng Yang +1
Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learning has been widely studied…
Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation
Zhijing Yang, Haocheng Lin, Zhihua Xu +4
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamental challenge in automated spa…
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
Zhenxuan Lu, Zhihua Xu, Zhijing Yang +4
Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth mov…
Learning Semantic-Aware Threshold for Multi-Label Image Recognition with Partial Labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang +4
Multi-label image recognition with partial labels (MLR-PL) is designed to train models using a mix of known and unknown labels. Traditional methods rely on semantic or feature corr…
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
Zhihua Xu, Tianshui Chen, Zhijing Yang +3
The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the tem…
Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang +3
Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since col…