5 papers
Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
Hongyi Fang, Chuwen Xie, Benjia Zhou +6
Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, t…
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
Chenggong Hu, Shaoyin Ma, Yi Wang +3
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability an…
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
Chenggong Hu, Yi Wang, Mengqi Xue +3
Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated T…
HuggingR: A Progressive Reasoning Framework for Discovering Optimal Model Companions
Shaoyin Ma, Chenggong Hu, Huiqiong Wang +3
Building effective LLM agents increasingly requires selecting appropriate AI models as tools from large open repositories (e.g., HuggingFace with > 2M models) based on natural lang…
Diffusion Model Quantization: A Review
Qian Zeng, Chenggong Hu, Mingli Song +1
Recent success of large text-to-image models has empirically underscored the exceptional performance of diffusion models in generative tasks. To facilitate their efficient deployme…