42 citations · 161 across the 34 of their papers we have counts for
6 papers · 1 filter
PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
Ahmadreza Jeddi, Hakki Can Karaimer, Hue Nguyen +8
RL post-training with verifiable rewards (RLVR) has become a practical route to eliciting chain-of-thought reasoning in vision--language models (VLMs), but scaling it in the visual…
Generative Point Tracking with Flow Matching
Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis +2
Tracking a point through a video can be a challenging task due to uncertainty arising from visual obfuscations, such as appearance changes and occlusions. Although current state-of…
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
Fereshteh Forghani, Jason J. Yu, Tristan Aumentado-Armstrong +2
Conventional depth-free multi-view datasets are captured using a moving monocular camera without metric calibration. The scales of camera positions in this monocular setting are am…
Revisiting Image Fusion for Multi-Illuminant White-Balance Correction
David Serrano-Lozano, Aditya Arora, Luis Herranz +3
White balance (WB) correction in scenes with multiple illuminants remains a persistent challenge in computer vision. Recent methods explored fusion-based approaches, where a neural…
Geometry-Aware Diffusion Models for Multiview Scene Inpainting
Ahmad Salimi, Tristan Aumentado-Armstrong, Marcus A. Brubaker +1
In this paper, we focus on 3D scene inpainting, where parts of an input image set, captured from different viewpoints, are masked out. The main challenge lies in generating plausib…
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
Harrish Thasarathan, Julian Forsyth, Thomas Fel +2
We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing…