222 citations · 356 across the 48 of their papers we have counts for
5 papers · 2 filters
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
Song Zhang, Yanlong Chen, Yilin Li +4
Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radically different visual evidence…
The Eleventh NTIRE 2026 Efficient Super-Resolution Challenge Report
Bin Ren, Hang Guo, Yan Shu +60
This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a…
ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration
Leheng Zhang, Wei Long, Yawei Li +3
Recently, Transformers have gained significant popularity in image restoration tasks such as image super-resolution and denoising, owing to their superior performance. However, bal…
Efficient Autoregressive Video Diffusion with Dummy Head
Hang Guo, Zhaoyang Jia, Jiahao Li +5
The autoregressive video diffusion model has recently gained considerable research interest due to its causal modeling and iterative denoising. In this work, we identify that the m…
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Yanlong Chen, Amirhossein Habibian, Luca Benini +1
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its pot…