activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

Prannay Kaul, Zhizhong Li, Hao Yang +4

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which…

cs.CV2024

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Hao Li, Shamit Lal, Zhiheng Li +9

We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…

cs.CV2024

NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality

Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7

We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…

cs.CV2024

Diffusion Soup: Model Merging for Text-to-Image Diffusion Models

Benjamin Biggs, Arjun Seshadri, Yang Zou +6

We present Diffusion Soup, a compartmentalization method for Text-to-Image Generation that averages the weights of diffusion models trained on sharded data. By construction, our ap…

cs.CV2024

Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model

Xiaolong Li, Jiawei Mo, Ying Wang +7

In this paper, we propose an effective two-stage approach named Grounded-Dreamer to generate 3D assets that can accurately follow complex, compositional text prompts while achievin…

cs.CV2024

Mixed-Query Transformer: A Unified Image Segmentation Architecture

Pei Wang, Zhaowei Cai, Hao Yang +3

Existing unified image segmentation models either employ a unified architecture across multiple tasks but use separate weights tailored to each dataset, or apply a single set of we…