auto-configuration 1diffusion models 1gpu memory management 1performance planning 1serving optimization 1
From the 1 of 8 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
Xuan Ai, Qingqing Yang, Peng Wang +4
Long-context inference in Large Language Models (LLMs) is bottlenecked by the quadratic computation complexity of attention and the substantial memory footprint of Key-Value (KV) c…
cs.CL2024
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
Shengnan Wang, Youhui Bai, Lin Zhang +7
Length generalization failure problem, namely the large language model (LLM) fails to generalize to texts longer than its maximum training length, greatly restricts the application…