Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning
Weihang Pan, Zhengxu Yu, Yuxiang Zhang +5
Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoni…
cs.LG2026
From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers
Binbin Lin, Wei Chen, Yalun Li +3
Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over th…
cs.LG2026
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
Ke Li, Dong An, Xiaoling Zang +6
Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activations contain outliers, but tha…