4 citations · 5 across the 3 of their papers we have counts for
8 papers · 1 filter
PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
Ahmadreza Jeddi, Hakki Can Karaimer, Hue Nguyen +8
RL post-training with verifiable rewards (RLVR) has become a practical route to eliciting chain-of-thought reasoning in vision--language models (VLMs), but scaling it in the visual…
Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
Weiming Ren, Raghav Goyal, Zhiming Hu +3
Generative super-resolution (GSR) currently sets the state-of-the-art in terms of perceptual image quality, overcoming the "regression-to-the-mean" blur of prior non-generative mod…
Extending Video Masked Autoencoders to 128 frames
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal +8
Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked…
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
Raghav Goyal, Wan-Cyuan Fan, Mennatullah Siam +1
Video Object Segmentation (VOS) has emerged as an increasingly important problem with availability of larger datasets and more complex and realistic settings, which involve long vi…
UniT: Unified Knowledge Transfer for Any-shot Object Detection and Segmentation
Siddhesh Khandelwal, Raghav Goyal, Leonid Sigal
Methods for object detection and segmentation rely on large scale instance-level annotations for training, which are difficult and time-consuming to collect. Efforts to alleviate t…
Ensemble based discriminative models for Visual Dialog Challenge 2018
Shubham Agarwal, Raghav Goyal
This manuscript describes our approach for the Visual Dialog Challenge 2018. We use an ensemble of three discriminative models with different encoders and decoders for our final su…