Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Base of RoPE Bounds Context Length
Xin Men, Mingyu Xu, Bingning Wang +4
Position embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation…
cs.CL2024
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Xin Men, Mingyu Xu, Qingyu Zhang +5
As Large Language Models (LLMs) continue to advance in performance, their size has escalated significantly, with current LLMs containing billions or even trillions of parameters. H…