Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Hao Li, Shamit Lal, Zhiheng Li +9
We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…
cs.CV2024
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6
Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…
cs.CV2024
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7
We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…