Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
Keqi Deng, Shaoshi Ling, Ruchao Fan +1
Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse attention alleviates this by lo…
cs.CL2025
Customizing Speech Recognition Model with Large Language Model Feedback
Shaoshi Ling, Guoli Ye
Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities…