94 citations · 116 across the 7 of their papers we have counts for
7 papers · 1 filter
MTet: Multi-domain Translation for English and Vietnamese
Chinh Ngo, Trieu H. Trinh, Long Phan +5
We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain…
ViT5: Pretrained Text-to-Text Transformer for Vietnamese Language Generation
Long Phan, Hieu Tran, Hieu Nguyen +1
We present ViT5, a pretrained Transformer-based encoder-decoder model for the Vietnamese language. With T5-style self-supervised pretraining, ViT5 is trained on a large corpus of h…
VieSum: How Robust Are Transformer-based Models on Vietnamese Summarization?
Hieu Nguyen, Long Phan, James Anibal +2
Text summarization is a challenging task within natural language processing that involves text generation from lengthy input sequences. While this task has been widely studied in E…
SPBERT: An Efficient Pre-training BERT on SPARQL Queries for Question Answering over Knowledge Graphs
Hieu Tran, Long Phan, James Anibal +2
In this paper, we propose SPBERT, a transformer-based language model pre-trained on massive SPARQL query logs. By incorporating masked language modeling objectives and the word str…
SciFive: a text-to-text transformer model for biomedical literature
Long N. Phan, James T. Anibal, Hieu Tran +4
In this report, we introduce SciFive, a domain-specific T5 model that has been pre-trained on large biomedical corpora. Our model outperforms the current SOTA methods (i.e. BERT, B…
Hierarchical Transformer Encoders for Vietnamese Spelling Correction
Hieu Tran, Cuong V. Dinh, Long Phan +1
In this paper, we propose a Hierarchical Transformer model for Vietnamese spelling correction problem. The model consists of multiple Transformer encoders and utilizes both charact…