activity
20182023
most citedOpen Vocabulary Object Detection with Proposal Mining and Prediction Equalization

7 citations · 26 across the 16 of their papers we have counts for

collaborators
Showing 2022Show all

12 papers · 1 filter

cs.CV2022

TaCo: Textual Attribute Recognition via Contrastive Learning

Chang Nie, Yiqing Hu, Yanqiu Qu +3

As textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing ap…

cs.CL2022

GMN: Generative Multi-modal Network for Practical Document Information Extraction

Haoyu Cao, Jiefeng Ma, Antai Guo +5

Document Information Extraction (DIE) has attracted increasing attention due to its various advanced applications in the real world. Although recent literature has already achieved…

cs.CV2022

OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification

Ye Liu, Lingfeng Qiao, Di Yin +4

Scene segmentation and classification (SSC) serve as a critical step towards the field of video structuring analysis. Intuitively, jointly learning of these two tasks can promote e…

cs.CV2022★ 3 cited

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai +3

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowl…

cs.CL2022★ 1 cited

RAAT: Relation-Augmented Attention Transformer for Relation Modeling in Document-Level Event Extraction

Yuan Liang, Zhuoxuan Jiang, Di Yin +1

In document-level event extraction (DEE) task, event arguments always scatter across sentences (across-sentence issue) and multiple events may lie in one document (multi-event issu…

cs.CV2022★ 1 cited

Contrastive Graph Multimodal Model for Text Classification in Videos

Ye Liu, Changchong Lu, Chen Lin +2

The extraction of text information in videos serves as a critical step towards semantic understanding of videos. It usually involved in two steps: (1) text recognition and (2) text…