3 papers
cs.CL2026
Correlation-Aware Structured Pruning for Large Language Models
Sicheng Xu, Hao Shi, Wei Zhang +6
Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods…
cs.LG2026
PRQuant: Permutation Residual Quantization for Low-Overhead Inference
Peiran Wang, Anqi Wang, Jiaying Zhao +9
Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotation, or residual-based approa…
cs.LG2025
Improving Monte Carlo Tree Search for Symbolic Regression
Zhengyao Huang, Daniel Zhengyu Huang, Tiannan Xiao +4
Symbolic regression aims to discover concise, interpretable mathematical expressions that satisfy desired objectives, such as fitting data, posing a highly combinatorial optimizati…