16 papers
InsHuman: Towards Natural and Identity-Preserving Human Insertion
Jie Li, Shulian Zhang, Yangyang Gao +4
Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure case…
Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations
Yuan Zhang, Chenyi Li, Guoqing Ma +8
Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often…
Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
Zheng Chen, Xun Zhang, Wenbo Li +7
The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed…
QuantVSR: Low-Bit Post-Training Quantization for Real-World Video Super-Resolution
Bowen Chai, Zheng Chen, Libo Zhu +3
Diffusion models have shown superior performance in real-world video super-resolution (VSR). However, the slow processing speeds and heavy resource consumption of diffusion models…
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
Kai Liu, Leyang Chen, Wenbo Li +5
Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. However, evaluations of unified multimodal models (UMMs) rem…
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
Shulian Zhang, Yong Guo, Long Peng +6
Video Face Enhancement (VFE) aims to restore high-quality facial regions from degraded video sequences, enabling a wide range of practical applications. Despite substantial progres…