papers

Publications (85)

cs.CL2021

Adaptive Nearest Neighbor Machine Translation

Xin Zheng, Zhirui Zhang, Junliang Guo +4

kNN-MT, recently proposed by Khandelwal et al. (2020a), successfully combines pre-trained neural machine translation (NMT) model with token-level k-nearest-neighbor (kNN) retrieval…

cs.CL2025

ReGLA: Refining Gated Linear Attention

Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh +2

Recent advancements in Large Language Models (LLMs) have set themselves apart with their exceptional performance in complex language modelling tasks. However, these models are also…

cs.CL2024

AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs

Alireza Ghaffari, Sharareh Younesian, Vahid Partovi Nia +2

The ever-growing computational complexity of Large Language Models (LLMs) necessitates efficient deployment strategies. The current state-of-the-art approaches for Post-training Qu…

cs.CL2021

QEMind: Alibaba's Submission to the WMT21 Quality Estimation Shared Task

Jiayi Wang, Ke Wang, Boxing Chen +3

Quality Estimation, as a crucial step of quality control for machine translation, has been explored for years. The goal is to investigate automatic methods for estimating the quali…

cs.LG2025

SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers

Aref Jafari, Yuhe Fan, Benyamin Jamialahmadi +3

Transformers have demonstrated strong performance across a wide range of sequence modeling tasks, but their quadratic attention complexity limits scalability to long sequences. Lin…

cs.LG2025

Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach

Alireza Ghaffari, Sharareh Younesian, Boxing Chen +2

As Large Language Models (LLMs) become increasingly computationally complex, developing efficient deployment strategies, such as quantization, becomes crucial. State-of-the-art Pos…