activity
20162022
most citedUnified Pretraining Framework for Document Understanding

17 citations · 24 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CV2022

SceneComposer: Any-Level Semantic Image Synthesis

Yu Zeng, Zhe Lin, Jianming Zhang +4

We propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More s…

cs.CV2022

Improving the Reliability for Confidence Estimation

Haoxuan Qu, Yanchao Li, Lin Geng Foo +3

Confidence estimation, a task that aims to evaluate the trustworthiness of the model's prediction output during deployment, has received lots of research attention recently, due to…

cs.CL202217 cited

Unified Pretraining Framework for Document Understanding

Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5

Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…

cs.CV2022

GALA: Toward Geometry-and-Lighting-Aware Object Search for Compositing

Sijie Zhu, Zhe Lin, Scott Cohen +3

Compositing-aware object search aims to find the most compatible objects for compositing given a background image and a query bounding box. Previous works focus on learning compati…

cs.CV2021

Multi-Scale Aligned Distillation for Low-Resolution Detection

Lu Qi, Jason Kuen, Jiuxiang Gu +5

In instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the…

cs.CV20217 cited

SelfDoc: Self-Supervised Document Representation Learning

Peizhao Li, Jiuxiang Gu, Jason Kuen +5

We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework…