activity
20192022
most citedTexSmart: A Text Understanding System for Fine-Grained NER and Enhanced Semantic Analysis

20 citations · 37 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL20229 cited

CO2Sum:Contrastive Learning for Factual-Consistent Abstractive Summarization

Wei Liu, Huanqin Wu, Wenjing Mu +3

Generating factual-consistent summaries is a challenging task for abstractive summarization. Previous works mainly encode factual information or perform post-correct/rank after dec…

cs.CL20211 cited

UniKeyphrase: A Unified Extraction and Generation Framework for Keyphrase Prediction

Huanqin Wu, Wei Liu, Lei Li +4

Keyphrase Prediction (KP) task aims at predicting several keyphrases that can summarize the main idea of the given document. Mainstream KP methods can be categorized into purely ge…

cs.CL20217 cited

LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding

Te-Lin Wu, Cheng Li, Mingyang Zhang +3

Document layout comprises both structural and visual (eg. font-sizes) information that is vital but often ignored by machine learning models. The few existing models which do use l…

cs.CL202020 cited

TexSmart: A Text Understanding System for Fine-Grained NER and Enhanced Semantic Analysis

Haisong Zhang, Lemao Liu, Haiyun Jiang +14

This technique report introduces TexSmart, a text understanding system that supports fine-grained named entity recognition (NER) and enhanced semantic analysis functionalities. Com…

cs.CL2019

CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases

Tao Yu, Rui Zhang, He Yang Er +21

We present CoSQL, a corpus for building cross-domain, general-purpose database (DB) querying dialogue systems. It consists of 30k+ turns plus 10k+ annotated SQL queries, obtained f…

cs.CL2019

SParC: Cross-Domain Semantic Parsing in Context

Tao Yu, Rui Zhang, Michihiro Yasunaga +16

We present SParC, a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences (12k+ individual questions annotated with SQL queries). It…