2 papers
cs.LG2026
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
Ruijie Miao, Zhiming Wang, Wang Li +4
Key-value (KV) caching is widely used to accelerate transformer inference, but its memory cost grows linearly with input length, limiting long-context deployment. Existing token ev…
cs.CL2023
N-gram Boosting: Improving Contextual Biasing with Normalized N-gram Targets
Wang Yau Li, Shreekantha Nadig, Karol Chang +5
Accurate transcription of proper names and technical terms is particularly important in speech-to-text applications for business conversations. These words, which are essential to…