3 papers
cs.LG2026
Towards Customized Multimodal Role-Play
Chao Tang, Jianzong Wu, Qingyu Shi +5
Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while…
cs.CV2025
An Empirical Study of GPT-4o Image Generation Capabilities
Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16
The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to brid…
cs.CV2025
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Jianzong Wu, Chao Tang, Jingbo Wang +3
Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-to-image generation models. However, these models often lack effe…