2 citations · 3 across the 6 of their papers we have counts for
6 papers
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
Rujiao Long, Pengfei Wang, Zhibo Yang +1
End-to-end visual information extraction (VIE) aims at integrating the hierarchical subtasks of VIE, including text spotting, word grouping, and entity labeling, into a unified fra…
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
Jianqiang Wan, Sibo Song, Wenwen Yu +6
Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Gene…
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang +4
Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations o…
Efficient Monaural Speech Enhancement using Spectrum Attention Fusion
Jinyu Long, Jetic Gū, Binhao Bai +3
Speech enhancement is a demanding task in automated speech processing pipelines, focusing on separating clean speech from noisy channels. Transformer based models have recently bes…
Modeling Entities as Semantic Points for Visual Information Extraction in the Wild
Zhibo Yang, Rujiao Long, Pengfei Wang +5
Recently, Visual Information Extraction (VIE) has been becoming increasingly important in both the academia and industry, due to the wide range of real-world applications. Previous…
Target-absent Human Attention
Zhibo Yang, Sounak Mondal, Seoyoung Ahn +3
The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user's attention. Computer vision models have been develope…