activity
20132023
most citedMASS: Masked Sequence to Sequence Pre-training for Language Generation

578 citations · 2.7k across the 97 of their papers we have counts for

collaborators
Showing cs.CLShow all

39 papers · 1 filter

cs.CL2023

Extract and Attend: Improving Entity Translation in Neural Machine Translation

Zixin Zeng, Rui Wang, Yichong Leng +4

While Neural Machine Translation(NMT) has achieved great progress in recent years, it still suffers from inaccurate translation of entities (e.g., person/organization name, locatio…

cs.CL2023★ 8 cited

MolXPT: Wrapping Molecules with Text for Generative Pre-training

Zequn Liu, Wei Zhang, Yingce Xia +5

Generative pre-trained Transformer (GPT) has demonstrates its great success in natural language processing and related techniques have been adapted into molecular modeling. Conside…

cs.CL2022

SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition

Yichong Leng, Xu Tan, Wenjie Liu +6

Error correction in automatic speech recognition (ASR) aims to correct those incorrect words in sentences generated by ASR models. Since recent ASR models usually have low word err…

cs.CL2021★ 11 cited

Pre-training Co-evolutionary Protein Representation via A Pairwise Masked Language Model

Liang He, Shizhuo Zhang, Lijun Wu +10

Understanding protein sequences is vital and urgent for biology, healthcare, and medicine. Labeling approaches are expensive yet time-consuming, while the amount of unlabeled data…

cs.CL2021

Discovering Drug-Target Interaction Knowledge from Biomedical Literature

Yutai Hou, Yingce Xia, Lijun Wu +6

The Interaction between Drugs and Targets (DTI) in human body plays a crucial role in biomedical science and applications. As millions of papers come out every year in the biomedic…

cs.CL2021★ 3 cited

A Survey on Low-Resource Neural Machine Translation

Rui Wang, Xu Tan, Renqian Luo +2

Neural approaches have achieved state-of-the-art accuracy on machine translation but suffer from the high cost of collecting large scale parallel data. Thus, a lot of research has…