34 papers
A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules
Abdul Mueez, Yogesh S. Rawat, Shruti Vyas
This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models on images where multiple defe…
Forget, Anticipate and Adapt: Test Time Training for Long Videos
Rajat Modi, Sebastian Noel, Xin Liang +1
Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during infere…
Asynchronous Perception Machine For Efficient Test-Time-Training
Rajat Modi, Yogesh Singh Rawat
In this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM can process patches of an image one at a…
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of s…
Learning to Deny: Action Denial in Multimodal Large Language Models
Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat
Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal +2
The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis R…