activity
20192022
most citedRead Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition

21 citations · 62 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CV2022

Intra-class Adaptive Augmentation with Neighbor Correction for Deep Metric Learning

Zheren Fu, Zhendong Mao, Bo Hu +2

Deep metric learning aims to learn an embedding space, where semantically similar samples are close together and dissimilar ones are repelled against. To explore more hard and info…

cs.CL20224 cited

UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction

Wei Tang, Benfeng Xu, Yuyue Zhao +4

Relational triple extraction is challenging for its difficulty in capturing rich correlations between entities and relations. Existing works suffer from 1) heterogeneous representa…

cs.CL20222 cited

Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity

Jiahao Li, Quan Wang, Zhendong Mao +3

Chinese spelling check (CSC) is a fundamental NLP task that detects and corrects spelling errors in Chinese texts. As most of these spelling errors are caused by phonetic similarit…

cs.CV202121 cited

Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition

Shancheng Fang, Hongtao Xie, Yuxin Wang +2

Linguistic knowledge is of great benefit to scene text recognition. However, how to effectively model linguistic rules in end-to-end deep networks remains a research challenge. In…

cs.CL20219 cited

Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation Extraction

Benfeng Xu, Quan Wang, Yajuan Lyu +2

Entities, as the essential elements in relation extraction tasks, exhibit certain structure. In this work, we formulate such structure as distinctive dependencies between mention p…

cs.CV20215 cited

Image Captioning with Context-Aware Auxiliary Guidance

Zeliang Song, Xiaofei Zhou, Zhendong Mao +1

Image captioning is a challenging computer vision task, which aims to generate a natural language description of an image. Most recent researches follow the encoder-decoder framewo…