activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

Prannay Kaul, Zhizhong Li, Hao Yang +4

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which…

cs.CV2024

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Hao Li, Shamit Lal, Zhiheng Li +9

We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…

cs.CV2024

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6

Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…

cs.CV2024

NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality

Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7

We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…

cs.CV2024

Diffusion Soup: Model Merging for Text-to-Image Diffusion Models

Benjamin Biggs, Arjun Seshadri, Yang Zou +6

We present Diffusion Soup, a compartmentalization method for Text-to-Image Generation that averages the weights of diffusion models trained on sharded data. By construction, our ap…

cs.CV2024

Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model

Xiaolong Li, Jiawei Mo, Ying Wang +7

In this paper, we propose an effective two-stage approach named Grounded-Dreamer to generate 3D assets that can accurately follow complex, compositional text prompts while achievin…