42 citations · 163 across the 36 of their papers we have counts for
5 papers · 1 filter
What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models
Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3
Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge acros…
Why Low-Light Cameras Go Color Blind: Removing Color Bias in Raw Denoising
Mohammad Mohammadi, Sina Honari, Stavros Tsogkas +6
Raw images inherently suffer from noise due to the stochastic nature of light and sensor hardware imperfections. As real photon counts fall, the ratio of this noise to the signal d…
AVIS: Adaptive Test-Time Scaling for Vision-Language Models
Ahmadreza Jeddi, Minh Ngoc Le, Amirhossein Kazerouni +8
Modern Vision-Language Models (VLMs) benefit from chain-of-thought prompting and test-time scaling, but these gains often come with prohibitive inference cost due to large visual c…
BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors
Dong Huo, Tristan Aumentado-Armstrong, Samrudhdhi B. Rangrej +8
Burst image super resolution (BISR) aims to construct a single high-resolution (HR) image by aggregating information from multiple low-resolution (LR) frames, relying on temporal r…
Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration
Amirhossein Kazerouni, Maitreya Suin, Tristan Aumentado-Armstrong +6
Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models (Ref-FR). However, such methods…