long-context memory 1model efficiency 1retrieval-augmented generation 1self-distillation 1transformer depth division 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
Hanzuo Liu, Xuan Qi, Chunyu Liu +6
The paper proposes CoMem, a method that stores intermediate transformer layer states as memory to enable efficient long‑context retrieval, showing that using lower‑mid layers for c…
cs.LG2026
SparseForge: Efficient Semi-Structured LLM Sparsification via Annealing of Hessian-Guided Soft-Mask
Liu Hanzuo, Chaofan Lin, Weixuan Sun +4
Semi-structured sparsity provides a practical path to accelerate large language models (LLMs) with native hardware support, but post-training semi-structured pruning often suffers…
cs.LG2024
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
Ao Shen, Zhiyao Li, Mingyu Gao
Serving numerous users and requests concurrently requires good fairness in Large Language Models (LLMs) serving system. This ensures that, at the same cost, the system can meet the…