73 citations · 90 across the 32 of their papers we have counts for
59 papers · 1 filter
A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules
Abdul Mueez, Yogesh S. Rawat, Shruti Vyas
This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models on images where multiple defe…
Forget, Anticipate and Adapt: Test Time Training for Long Videos
Rajat Modi, Sebastian Noel, Xin Liang +1
Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during infere…
Learning to Deny: Action Denial in Multimodal Large Language Models
Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat
Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal +2
The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis R…
MolSight: Molecular Property Prediction with Images
Aaditya Baranwal, Akshaj Gupta, Yogesh S Rawat +1
Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has received less focus in favour…
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5
Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity t…