collaborators

7 papers

cs.CV2026

InsHuman: Towards Natural and Identity-Preserving Human Insertion

Jie Li, Shulian Zhang, Yangyang Gao +4

Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure case…

cs.CV2026

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

Zheng Chen, Xun Zhang, Wenbo Li +7

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed…

cs.CV2025

PocketSR: The Super-Resolution Expert in Your Pocket Mobiles

Haoze Sun, Linfeng Jiang, Fan Li +9

Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging larg…

cs.CV2025

LoViC: Efficient Long Video Generation with Context Compression

Jiaxiu Jiang, Wenbo Li, Jingjing Ren +5

Despite recent advances in diffusion transformers (DiTs) for text-to-video generation, scaling to long-duration content remains challenging due to the quadratic complexity of self-…

cs.CV2025

PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement

ZhanFeng Feng, Long Peng, Xin Di +7

Multi-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which…

cs.CV2025

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

Jingjing Ren, Wenbo Li, Zhongdao Wang +9

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals. While diffusion transformers (DiTs) have demonstrated remarkable capabilities…