1 citations · 1 across the 7 of their papers we have counts for
4 papers · 1 filter
Beyond Hard Writes and Rigid Preservation: Soft Recursive Least-Squares for Lifelong LLM Editing
Xinyu Wang, Sicheng Lyu, Yu Gu +4
Model editing updates a pre-trained LLM with new facts or rules without retraining while preserving unrelated behavior. In real deployment, edits arrive as long streams, creating a…
OJBKQ: Objective-Joint Babai-Klein Quantization
Xinyu Wang, Ziyu Zhao, Peng Lu +2
Post-training quantization (PTQ) is widely used to compress large language models without retraining. However, many existing weight-only methods rely on heuristic objectives and gr…
Mamba Modulation: On the Length Generalization of Mamba
Peng Lu, Jerry Huang, Qiuhao Zeng +4
The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space…
Calibrated Language Models and How to Find Them with Label Smoothing
Jerry Huang, Peng Lu, Qiuhao Zeng
Recent advances in natural language processing (NLP) have opened up greater opportunities to enable fine-tuned large language models (LLMs) to behave as more powerful interactive a…