most citedExtrapolating Large Language Models to Non-English by Aligning Languages

8 citations · 13 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2023

Roles of Scaling and Instruction Tuning in Language Perception: Model vs. Human Attention

Changjiang Gao, Shujian Huang, Jixing Li +1

Recent large language models (LLMs) have revealed strong abilities to understand natural language. Since most of them share the same basic structure, i.e. the transformer block, po…

cs.CL2023

IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation Systems

Xu Huang, Zhirui Zhang, Ruize Gao +6

We present IMTLab, an open-source end-to-end interactive machine translation (IMT) system platform that enables researchers to quickly build IMT systems with state-of-the-art model…

cs.CL2023

Only 5\% Attention Is All You Need: Efficient Long-range Document-level Neural Machine Translation

Zihan Liu, Zewei Sun, Shanbo Cheng +2

Document-level Neural Machine Translation (DocNMT) has been proven crucial for handling discourse phenomena by introducing document-level context information. One of the most impor…

cs.CV2023

Food-500 Cap: A Fine-Grained Food Caption Benchmark for Evaluating Vision-Language Models

Zheng Ma, Mianzhi Pan, Wenhan Wu +4

Vision-language models (VLMs) have shown impressive performance in substantial downstream multi-modal tasks. However, only comparing the fine-tuned performance on downstream tasks…

cs.CL20238 cited

Extrapolating Large Language Models to Non-English by Aligning Languages

Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6

Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…

cs.CL2023

INK: Injecting kNN Knowledge in Nearest Neighbor Machine Translation

Wenhao Zhu, Jingjing Xu, Shujian Huang +2

Neural machine translation has achieved promising results on many translation tasks. However, previous studies have shown that neural models induce a non-smooth representation spac…