167 citations · 223 across the 18 of their papers we have counts for
26 papers · 1 filter
Patched Denoising Diffusion Models For High-Resolution Image Synthesis
Zheng Ding, Mengqi Zhang, Jiajun Wu +1
We propose an effective denoising diffusion model for generating high-resolution images (e.g., 1024512), trained on small-size image patches (e.g., 6464). We name o…
DocTr: Document Transformer for Structured Information Extraction in Documents
Haofu Liao, Aruni RoyChowdhury, Weijian Li +6
We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based for…
Distilling Large Vision-Language Model with Out-of-Distribution Generalizability
Xuanlin Li, Yunhao Fang, Minghua Liu +3
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sen…
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
Zhaoyang Zhang, Yantao Shen, Kunyu Shi +7
We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, result…
DiffusionRig: Learning Personalized Priors for Facial Appearance Editing
Zheng Ding, Xuaner Zhang, Zhihao Xia +3
We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person'…
Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction
Hansheng Chen, Jiatao Gu, Anpei Chen +4
3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a compreh…