5 papers
InsHuman: Towards Natural and Identity-Preserving Human Insertion
Jie Li, Shulian Zhang, Yangyang Gao +4
Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure case…
ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization
Ronghao Zhang, Shuaicheng Niu, Qi Deng +3
Test-time adaptation (TTA) aims to improve model robustness under distribution shifts by adapting to unlabeled test data, but most existing methods rely on backpropagation (BP), wh…
FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
Fuhan Cai, Yong Guo, Jie Li +3
Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, the…
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
Shulian Zhang, Yong Guo, Long Peng +6
Video Face Enhancement (VFE) aims to restore high-quality facial regions from degraded video sequences, enabling a wide range of practical applications. Despite substantial progres…
Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic Teacher
Yong Guo, Shulian Zhang, Haolin Pan +3
Knowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. We find…