4 papers · 1 filter
Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion
Zhen Cheng, Hao-Bo Yang, Wan-Yi Huang +1
Key-Value (KV) cache memory and bandwidth increasingly dominate large language model inference cost in long-context and long-generation regimes. Architectures such as multi-head la…
ILRe: Intermediate Layer Retrieval for Context Compression in Causal Language Models
Manlai Liang, Mandi Liu, Jiangzhou Ji +4
Large Language Models (LLMs) have demonstrated success across many benchmarks. However, they still exhibit limitations in long-context scenarios, primarily due to their short effec…
Lag-Relative Sparse Attention In Long Context Training
Manlai Liang, Wanyi Huang, Mandi Liu +2
Large Language Models (LLMs) have made significant strides in natural language processing and generation, yet their ability to handle long-context input remains constrained by the…
A Training-free LLM Framework with Interaction between Contextually Related Subtasks in Solving Complex Tasks
Hongjia Liu, Jinlong Li
Large language models (LLMs) have shown remarkable capabilities in solving complex tasks. Recent work has explored decomposing such tasks into subtasks with independent contexts. H…