3 papers
cs.LG2026
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
Xin Teng, Canyu Zhang, Shaoyi Zheng +3
The paper proposes InfoFlow KV, a method that uses an attention‑norm signal to identify which key‑value cache tokens should be recomputed during retrieval‑augmented generation, imp…
cs.CL2025
Submodular Context Partitioning and Compression for In-Context Learning
Shaoyi Zheng, Canyu Zhang, Tianyi Zhou +1
In-context learning (ICL) enables efficient few-shot learning in large language models (LLMs) without training, but suffers from the quadratic input complexity of transformers, lim…
cs.LG2025
Tilted Sharpness-Aware Minimization
Tian Li, Tianyi Zhou, Jeffrey A. Bilmes
Sharpness-Aware Minimization (SAM) has been demonstrated to improve the generalization performance of overparameterized models by seeking flat minima on the loss landscape through…