2 papers
cs.LG2026
Shape Mutating Expert Compression:LorExperts and BTExperts
Inesh Chakrabarti, Sourjya Roy, Bowen Bao +3
Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert p…
cs.LG2026
UltraQuant: 4-bit KV Caching for Context-Heavy Agents
Inesh Chakrabarti, David Limpus, Aditi Ghai Rana +4
Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrency determines whether the serving system c…