Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
SparseAccelerate: Efficient Long-Context Inference for Mid-Range GPUs
James Vo
As Large Language Models (LLMs) scale to longer context windows, the computational cost of attention mechanisms, which traditionally grows quadratically with input length, presents…
cs.CL2024
Combining Entropy and Matrix Nuclear Norm for Enhanced Evaluation of Language Models
James Vo
As large language models (LLMs) continue to advance, the need for precise and efficient evaluation metrics becomes more pressing. Traditional approaches, while informative, often f…
cs.CL2024
Transformer Layer Injection: A Novel Approach for Efficient Upscaling of Large Language Models
James Vo
In this paper, we propose Transformer Layer Injection (TLI), a novel method for efficiently upscaling large language models (LLMs) while minimizing computational costs and maintain…