2 papers
cs.CL2026
LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding
Haocheng Xia, Mihir Pamnani, Hanxi Fang +2
Key-value (KV) caching accelerates inference of large language models (LLMs) by reusing past computations for generated tokens. Its importance becomes even greater in long-context…
cs.LG2025
SIMU: Selective Influence Machine Unlearning
Anu Agarwal, Mihir Pamnani, Dilek Hakkani-Tur
The undesired memorization of sensitive information by Large Language Models (LLMs) has emphasized the need for safety mechanisms that can regulate model behavior. This has led to…