activity
20222024
most citedModeling Entities as Semantic Points for Visual Information Extraction in the Wild

2 citations · 3 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction

Rujiao Long, Pengfei Wang, Zhibo Yang +1

End-to-end visual information extraction (VIE) aims at integrating the hierarchical subtasks of VIE, including text spotting, word grouping, and entity labeling, into a unified fra…

cs.CV20241 cited

OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Jianqiang Wan, Sibo Song, Wenwen Yu +6

Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Gene…

cs.CV2024

LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training

Rujiao Long, Hangdi Xing, Zhibo Yang +4

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations o…

cs.SD2023

Efficient Monaural Speech Enhancement using Spectrum Attention Fusion

Jinyu Long, Jetic Gū, Binhao Bai +3

Speech enhancement is a demanding task in automated speech processing pipelines, focusing on separating clean speech from noisy channels. Transformer based models have recently bes…

cs.CV20232 cited

Modeling Entities as Semantic Points for Visual Information Extraction in the Wild

Zhibo Yang, Rujiao Long, Pengfei Wang +5

Recently, Visual Information Extraction (VIE) has been becoming increasingly important in both the academia and industry, due to the wide range of real-world applications. Previous…

cs.CV2022

Target-absent Human Attention

Zhibo Yang, Sounak Mondal, Seoyoung Ahn +3

The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user's attention. Computer vision models have been develope…