works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.CV2026

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

Jian Zou, Xiaoyu Xu, Zhihua Wang +3

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token se…

cs.CV2026

LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models

Bowen Chen, Shreshth Saini, Balu Adsumilli +1

LumaGuide is a training‑free framework that steers the sampling of pretrained diffusion models by matching target luminance distributions, enabling high dynamic range (HDR) image g…

cs.AI2026

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

Shreshth Saini, Neil Birkbeck, Yilin Wang +2

Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching ma…

cs.CV2026

Context and Pixel Aware Large Language Model for Video Quality Assessment

Wen Wen, Yaohong Wu, Yue Sheng +3

Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-…

cs.CV2026

Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing

Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck +3

With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challenges. Static cropping and bord…

cs.CV2026

LumaFlux: Lifting 8-Bit Worlds to HDR Reality with Physically-Guided Diffusion Transformers

Shreshth Saini, Hakan Gedik, Neil Birkbeck +3

The rapid adoption of HDR-capable devices has created a pressing need to convert the 8-bit Standard Dynamic Range (SDR) content into perceptually and physically accurate 10-bit Hig…