3 papers
cs.LG2026
XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression
Jundong Hu, Shekar Ramachandran
Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictab…
cs.AI2026
The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents
Jundong Hu, Shekar Ramachandran
Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capabil…
cs.LG2026
The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
Jundong Hu, Shekar Ramachandran
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study w…