Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression
Jundong Hu, Shekar Ramachandran
Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictab…
cs.LG2026
The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
Jundong Hu, Shekar Ramachandran
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study w…