Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Yushi Bai, Qian Dong, Ting Jiang +5
Long-context agentic workflows have emerged as a defining use case for large language models, making attention efficiency critical for both inference speed and serving cost. Sparse…
cs.CL2025
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
Tingyu Jiang, Shen Li, Yiyao Song +6
Instruction tuning plays a critical role in enhancing the performance and efficiency of Large Language Models (LLMs). Its success depends not only on the quality of the instruction…