7 papers · 1 filter
See then Tell: Enhancing Key Information Extraction with Vision Grounding
Shuhang Liu, Zhenrong Zhang, Pengfei Hu +5
In the digital era, the ability to understand visually rich documents that integrate text, complex layouts, and imagery is critical. Traditional Key Information Extraction (KIE) me…
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
Hanbo Cheng, Limin Lin, Chenyu Liu +5
Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diff…
RFL: Simplifying Chemical Structure Recognition with Ring-Free Language
Qikai Chang, Mingjun Chen, Changpeng Pi +6
The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional s…
Skeleton and Font Generation Network for Zero-shot Chinese Character Generation
Mobai Xue, Jun Du, Zhenrong Zhang +5
Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
Haotian Wang, Yuzhe Weng, Yueyan Li +10
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In t…
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
Zhenrong Zhang, Shuhang Liu, Pengfei Hu +4
In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual…