3 papers
cs.CV2024
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
Jianyi Zhang, Yufan Zhou, Jiuxiang Gu +5
Diffusion models have demonstrated exceptional capabilities in generating a broad spectrum of visual content, yet their proficiency in rendering text is still limited: they often g…
cs.CV2024
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
Curtis Wigington
Document capture applications on smartphones have emerged as popular tools for digitizing documents. For many individuals, capturing documents with their smartphones is more conven…
cs.CV2024
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
Sanket Biswas, Rajiv Jain, Vlad I. Morariu +5
While the generation of document layouts has been extensively explored, comprehensive document generation encompassing both layout and content presents a more complex challenge. Th…