Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Value-Aware Stochastic KV Cache Eviction for Reasoning Models
Ting-Yun Chang, Harvey Yiyun Fu, Deqing Fu +3
Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction methods reduce this cost by e…
cs.LG2025
Why Do Some Inputs Break Low-Bit LLM Quantization?
Ting-Yun Chang, Muru Zhang, Jesse Thomason +1
Low-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples. We analyze diverse 3-4…