works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

DynEval: Holistic Evaluations of T2I Generative Models in the Wild

Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3

The paper introduces DynEval, a dynamic evaluation framework that jointly assesses text-to-image alignment and image quality for T2I models, using large synthetic datasets and a di…

cs.CV2025

O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model

Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit +2

While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, th…

cs.CV2025

FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation

M Yashwanth, Sampath Koti, Arunabh Singh +2

We address the Federated source-Free Domain Adaptation (FFreeDA) problem, with clients holding unlabeled data with significant inter-client domain gaps. The FFreeDA setup constrain…

cs.CV2025

TTRV: Test-Time Reinforcement Learning for Vision Language Models

Akshit Singh, Shyam Marjit, Wei Lin +7

Existing methods for extracting reward signals in Reinforcement Learning typically rely on labeled data and dedicated training splits, a setup that contrasts with how humans learn…

cs.CV2025

CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives

Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri +1

With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transfo…

cs.CV2025

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models

Priyank Pathak, Shyam Marjit, Shruti Vyas +1

Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. Howev…