activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

Jiamian Wang, Ruiyi Zhang, Tong Yu +5

Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples serve as the training envir…

cs.CV2026

Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis

Runzhou Liu, Hailey Weingord, Sejal Mittal +18

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important…

cs.CV2026

Localizing Knowledge in Diffusion Transformers

Arman Zarei, Samyadeep Basu, Keivan Rezaei +3

Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has e…

cs.CV2025

SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control

Arman Zarei, Samyadeep Basu, Mobina Pournemat +3

Instruction-based image editing models have recently achieved impressive performance, enabling complex edits to an input image from a multi-instruction prompt. However, these model…

cs.CV2025

Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings

Arman Zarei, Keivan Rezaei, Samyadeep Basu +4

Text-to-image diffusion-based generative models have the stunning ability to generate photo-realistic images and achieve state-of-the-art low FID scores on challenging image genera…

cs.CV2024

Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP

Sriram Balasubramanian, Samyadeep Basu, Soheil Feizi

Recent work has explored how individual components of the CLIP-ViT model contribute to the final representation by leveraging the shared image-text representation space of CLIP. Th…