activity
20152023
most citedTraining Deeper Convolutional Networks with Deep Supervision

167 citations · 223 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV20233 cited

Patched Denoising Diffusion Models For High-Resolution Image Synthesis

Zheng Ding, Mengqi Zhang, Jiajun Wu +1

We propose an effective denoising diffusion model for generating high-resolution images (e.g., 1024512), trained on small-size image patches (e.g., 6464). We name o…

cs.CV20234 cited

DocTr: Document Transformer for Structured Information Extraction in Documents

Haofu Liao, Aruni RoyChowdhury, Weijian Li +6

We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based for…

cs.CV2023

Distilling Large Vision-Language Model with Out-of-Distribution Generalizability

Xuanlin Li, Yunhao Fang, Minghua Liu +3

Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sen…

cs.CV2023

Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts

Zhaoyang Zhang, Yantao Shen, Kunyu Shi +7

We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, result…

cs.CV2023

DiffusionRig: Learning Personalized Priors for Facial Appearance Editing

Zheng Ding, Xuaner Zhang, Zhihao Xia +3

We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person'…

cs.CV2023

Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction

Hansheng Chen, Jiatao Gu, Anpei Chen +4

3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a compreh…