20 papers · 1 filter
DiffCVE: Diffusion-based Compressed Video Enhancement
Wenqiang Xiao, Wenzhuo Ma, Junxi Zhang +1
Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Recent diffusion models have dem…
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
Yiling Gao, Hongchen Wei, Zhenzhong Chen
In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To a…
The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview
Jiatong Li, Zheng Chen, Kai Liu +91
This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the resulting outcomes. The challenge…
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
Wenzhuo Ma, Zhenzhong Chen
The practical deployment of diffusion-based Neural Video Compression (NVC) faces critical challenges, including severe information loss, prohibitive inference latency, and poor tem…
SemCo: Toward Semantic Coherent Visual Relationship Forecasting
Yangjun Ou, Yao Liu, Li Mi +1
Visual Relationship Forecasting (VRF) aims to anticipate relations among objects without observing future visual content. The task relies on capturing and modeling the semantic coh…
SemPT: Semantic Prompt Tuning for Vision-Language Models
Xiao Shi, Yangjun Ou, Zhenzhong Chen
Visual transfer learning for unseen categories presents an active research topic yet a challenging task, due to the inherent conflict between preserving category-specific represent…