activity
20182026
most citedIn Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

73 citations · 90 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

59 papers · 1 filter

cs.CV2026

A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules

Abdul Mueez, Yogesh S. Rawat, Shruti Vyas

This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models on images where multiple defe…

cs.CV2026

Forget, Anticipate and Adapt: Test Time Training for Long Videos

Rajat Modi, Sebastian Noel, Xin Liang +1

Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during infere…

cs.CV2026

Learning to Deny: Action Denial in Multimodal Large Language Models

Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat

Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…

cs.CV2026

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal +2

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis R…

cs.CV2026

MolSight: Molecular Property Prediction with Images

Aaditya Baranwal, Akshaj Gupta, Yogesh S Rawat +1

Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has received less focus in favour…

cs.CV2026

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity t…