164 citations · 916 across the 31 of their papers we have counts for
21 papers
Augmenting Language Models with Long-Term Memory
Weizhi Wang, Li Dong, Hao Cheng +4
Existing large language models (LLMs) can only afford fix-sized inputs due to the input length limit, preventing them from utilizing rich long-context information from past inputs.…
On the Off-Target Problem of Zero-Shot Multilingual Neural Machine Translation
Liang Chen, Shuming Ma, Dongdong Zhang +2
While multilingual neural machine translation has achieved great success, it suffers from the off-target issue, where the translation is in the wrong language. This problem is more…
Multiview Identifiers Enhanced Generative Retrieval
Yongqi Li, Nan Yang, Liang Wang +2
Instead of simply matching a query to pre-existing passages, generative retrieval generates identifier strings of passages as the retrieval target. At a cost, the identifier must b…
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Tianrui Wang, Long Zhou, Ziqiang Zhang +6
Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we propose V…
One-stop Training of Multiple Capacity Models
Lan Jiang, Haoyang Huang, Dongdong Zhang +2
Training models with varying capacities can be advantageous for deploying them in different scenarios. While high-capacity models offer better performance, low-capacity models requ…
Dual-Alignment Pre-training for Cross-lingual Sentence Embedding
Ziheng Li, Shaohan Huang, Zihan Zhang +7
Recent studies have shown that dual encoder models trained with the sentence-level translation ranking task are effective methods for cross-lingual sentence embedding. However, our…