activity
20152026
most citedPosition, Padding and Predictions: A Deeper Look at Position Information in CNNs

42 citations · 163 across the 36 of their papers we have counts for

collaborators
Showing 2026Show all

5 papers · 1 filter

cs.CV2026

What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3

Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge acros…

cs.CV2026

Why Low-Light Cameras Go Color Blind: Removing Color Bias in Raw Denoising

Mohammad Mohammadi, Sina Honari, Stavros Tsogkas +6

Raw images inherently suffer from noise due to the stochastic nature of light and sensor hardware imperfections. As real photon counts fall, the ratio of this noise to the signal d…

cs.CV2026

AVIS: Adaptive Test-Time Scaling for Vision-Language Models

Ahmadreza Jeddi, Minh Ngoc Le, Amirhossein Kazerouni +8

Modern Vision-Language Models (VLMs) benefit from chain-of-thought prompting and test-time scaling, but these gains often come with prohibitive inference cost due to large visual c…

cs.CV2026

BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors

Dong Huo, Tristan Aumentado-Armstrong, Samrudhdhi B. Rangrej +8

Burst image super resolution (BISR) aims to construct a single high-resolution (HR) image by aggregating information from multiple low-resolution (LR) frames, relying on temporal r…

cs.CV2026

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

Amirhossein Kazerouni, Maitreya Suin, Tristan Aumentado-Armstrong +6

Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models (Ref-FR). However, such methods…