most citedGTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation

25 citations · 76 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL20232 cited

MT4CrossOIE: Multi-stage Tuning for Cross-lingual Open Information Extraction

Tongliang Li, Zixiang Wang, Linzheng Chai +8

Cross-lingual open information extraction aims to extract structured information from raw text across multiple languages. Previous work uses a shared cross-lingual pre-trained mode…

cs.CV202314 cited

MIT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Lei Li, Yuwei Yin, Shicheng Li +9

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress i…

cs.CV20232 cited

TTIDA: Controllable Generative Data Augmentation via Text-to-Text and Text-to-Image Models

Yuwei Yin, Jean Kaddour, Xiang Zhang +4

Data augmentation has been established as an efficacious approach to supplement useful information for low-resource datasets. Traditional augmentation techniques such as noise inje…

cs.CL20239 cited

HanoiT: Enhancing Context-aware Translation via Selective Context

Jian Yang, Yuwei Yin, Shuming Ma +7

Context-aware neural machine translation aims to use the document-level context to improve translation quality. However, not all words in the context are helpful. The irrelevant or…

cs.CL202211 cited

UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine Translation

Jian Yang, Yuwei Yin, Shuming Ma +5

Most translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enable…

cs.CL202213 cited

HLT-MT: High-resource Language-specific Training for Multilingual Neural Machine Translation

Jian Yang, Yuwei Yin, Shuming Ma +3

Multilingual neural machine translation (MNMT) trained in multiple language pairs has attracted considerable attention due to fewer model parameters and lower training costs by sha…