6 papers
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
Yibo Peng, Peng Xia, Ding Zhong +6
Despite the rapid advancements in Multimodal Large Language Models (MLLMs), a critical question regarding their visual grounding mechanism remains unanswered: do these models genui…
Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style Transfer
Yanqi Ge, Jiaqi Liu, Qingnan Fan +6
In this work, we target the task of text-driven style transfer in the context of text-to-image (T2I) diffusion models. The main challenge is consistent structure preservation while…
Ovis-Image Technical Report
Guo-Hua Wang, Liangfu Cao, Tianyu Cui +8
We introduce , a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational c…
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
Jiaqi Liu, Tao Huang, Chang Xu
Recent advances in autoregressive (AR) models have demonstrated their potential to rival diffusion models in image synthesis. However, for complex spatially-conditioned generation,…
Trade-offs in Image Generation: How Do Different Dimensions Interact?
Sicheng Zhang, Binzhu Xie, Zhonghao Yan +7
Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, mo…
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
Jiaqi Liu, Jichao Zhang, Paolo Rota +1
The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS…