From the 1 of 6 linked papers with an AI index.
6 papers
DynEval: Holistic Evaluations of T2I Generative Models in the Wild
Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane +3
The paper introduces DynEval, a dynamic evaluation framework that jointly assesses text-to-image alignment and image quality for T2I models, using large synthetic datasets and a di…
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit +2
While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, th…
FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
M Yashwanth, Sampath Koti, Arunabh Singh +2
We address the Federated source-Free Domain Adaptation (FFreeDA) problem, with clients holding unlabeled data with significant inter-client domain gaps. The FFreeDA setup constrain…
TTRV: Test-Time Reinforcement Learning for Vision Language Models
Akshit Singh, Shyam Marjit, Wei Lin +7
Existing methods for extracting reward signals in Reinforcement Learning typically rely on labeled data and dedicated training splits, a setup that contrasts with how humans learn…
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri +1
With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transfo…
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
Priyank Pathak, Shyam Marjit, Shruti Vyas +1
Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. Howev…