Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
OJBKQ: Objective-Joint Babai-Klein Quantization
Xinyu Wang, Ziyu Zhao, Peng Lu +2
Post-training quantization (PTQ) is widely used to compress large language models without retraining. However, many existing weight-only methods rely on heuristic objectives and gr…
cs.LG2025
Mamba Modulation: On the Length Generalization of Mamba
Peng Lu, Jerry Huang, Qiuhao Zeng +4
The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space…
cs.LG2025
Calibrated Language Models and How to Find Them with Label Smoothing
Jerry Huang, Peng Lu, Qiuhao Zeng
Recent advances in natural language processing (NLP) have opened up greater opportunities to enable fine-tuned large language models (LLMs) to behave as more powerful interactive a…