4 papers · 1 filter
LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
Ying Shen, Zhiyang Xu, Jiuhai Chen +6
Recent advances in multimodal foundation models unifying image understanding and generation have opened exciting avenues for tackling a wide range of vision-language tasks within a…
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
Kaijie Chen, Zihao Lin, Zhiyang Xu +5
Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a w…
Holistic Evaluation for Interleaved Text-and-Image Generation
Minqian Liu, Zhiyang Xu, Zihao Lin +4
Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despit…
DECDM: Document Enhancement using Cycle-Consistent Diffusion Models
Jiaxin Zhang, Joy Rimchala, Lalla Mouatadid +2
The performance of optical character recognition (OCR) heavily relies on document image quality, which is crucial for automatic document processing and document intelligence. Howev…