Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Beyond KV Caching: Shared Attention for Efficient LLMs
Bingli Liao, Danilo Vasconcellos Vargas
The efficiency of large language models (LLMs) remains a critical challenge, particularly in contexts where computational resources are limited. Traditional attention mechanisms in…
cs.CL2024
Extending Token Computation for LLM Reasoning
Bingli Liao, Danilo Vasconcellos Vargas
Large Language Models (LLMs) are pivotal in advancing natural language processing but often struggle with complex reasoning tasks due to inefficient attention distributions. In thi…