1 citations · 1 across the 3 of their papers we have counts for
4 papers
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
Renjie Xie, Juncheng Yang, Aoting Hu +4
Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains…
HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory
Chengda Lu, Xiaoyu Fan, Wei Xu
While Large Language Models (LLMs) achieve strong performance across diverse tasks, their inference dynamics remain poorly understood because of the limited resolution of existing…
KVDirect: Distributed Disaggregated LLM Inference
Shiyang Chen, Rain Jiang, Dezhi Yu +6
Large Language Models (LLMs) have become the new foundation for many applications, reshaping human society like a storm. Disaggregated inference, which separates prefill and decode…
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
Yang Wu, Huayi Zhang, Yizheng Jiao +6
Instruction tuning has underscored the significant potential of large language models (LLMs) in producing more human controllable and effective outputs in various domains. In this…