164 citations · 544 across the 15 of their papers we have counts for
12 papers · 1 filter
BitNet: Scaling 1-bit Transformers for Large Language Models
Hongyu Wang, Shuming Ma, Li Dong +7
The increasing size of large language models has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. In this work, we int…
Retentive Network: A Successor to Transformer for Large Language Models
Yutao Sun, Li Dong, Shaohan Huang +5
In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and g…
LongNet: Scaling Transformers to 1,000,000,000 Tokens
Jiayu Ding, Shuming Ma, Li Dong +5
Scaling sequence length has become a critical demand in the era of large language models. However, existing methods struggle with either computational complexity or model expressiv…
Kosmos-2: Grounding Multimodal Large Language Models to the World
Zhiliang Peng, Wenhui Wang, Li Dong +4
We introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual…
On the Off-Target Problem of Zero-Shot Multilingual Neural Machine Translation
Liang Chen, Shuming Ma, Dongdong Zhang +2
While multilingual neural machine translation has achieved great success, it suffers from the off-target issue, where the translation is in the wrong language. This problem is more…
Discourse Centric Evaluation of Machine Translation with a Densely Annotated Parallel Corpus
Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma +3
Several recent papers claim human parity at sentence-level Machine Translation (MT), especially in high-resource languages. Thus, in response, the MT community has, in part, shifte…