activity
20242026
collaborators

34 papers

cs.CV2026

A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules

Abdul Mueez, Yogesh S. Rawat, Shruti Vyas

This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models on images where multiple defe…

cs.CV2026

Forget, Anticipate and Adapt: Test Time Training for Long Videos

Rajat Modi, Sebastian Noel, Xin Liang +1

Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during infere…

cs.CV2026

Asynchronous Perception Machine For Efficient Test-Time-Training

Rajat Modi, Yogesh Singh Rawat

In this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM can process patches of an image one at a…

cs.CV2026

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat

This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of s…

cs.CV2026

Learning to Deny: Action Denial in Multimodal Large Language Models

Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat

Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…

cs.CV2026

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal +2

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis R…