12 papers
Bounded-Compute Multimodal Regression for Product-Rating Prediction
William Leach, Ru He, Sizhuo Ma +4
Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generation and dynamic visual process…
Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
Jinpei Guo, Yifei Ji, Shengwei Wang +8
Diffusion models have recently shown promising results for video super-resolution (VSR). However, directly adapting generative diffusion models to VSR can result in redundancy, sin…
The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview
Zheng Chen, Kai Liu, Jingkai Wang +150
This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…
Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
Eric M. Chen, Di Liu, Sizhuo Ma +8
Despite the increasing popularity of avatar systems such as Snapchat Bitmojis, existing production avatar platforms face several limitations, such as a limited number of predefined…
Velocity Disambiguation for Video Frame Interpolation
Zhihang Zhong, Yiming Zhang, Wei Wang +5
Existing video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep t ("time indexing"), which struggles to predict precise object movement…
gQIR: Generative Quanta Image Reconstruction
Aryan Garg, Sizhuo Ma, Mohit Gupta
Capturing high-quality images from only a few detected photons is a fundamental challenge in computational imaging. Single-photon avalanche diode (SPAD) sensors promise high-qualit…