collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

DynEval: Holistic Evaluations of T2I Generative Models in the Wild

Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3

Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, espec…

cs.CV2025

O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model

Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit +2

While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, th…

cs.CV2025

FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation

M Yashwanth, Sampath Koti, Arunabh Singh +2

We address the Federated source-Free Domain Adaptation (FFreeDA) problem, with clients holding unlabeled data with significant inter-client domain gaps. The FFreeDA setup constrain…

cs.CV2025

TTRV: Test-Time Reinforcement Learning for Vision Language Models

Akshit Singh, Shyam Marjit, Wei Lin +7

Existing methods for extracting reward signals in Reinforcement Learning typically rely on labeled data and dedicated training splits, a setup that contrasts with how humans learn…

cs.CV2025

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models

Priyank Pathak, Shyam Marjit, Shruti Vyas +1

Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. Howev…

cs.CV2024

DiffuseKronA: A Parameter Efficient Fine-tuning Method for Personalized Diffusion Models

Shyam Marjit, Harshit Singh, Nityanand Mathur +3

In the realm of subject-driven text-to-image (T2I) generative models, recent developments like DreamBooth and BLIP-Diffusion have led to impressive results yet encounter limitation…