4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.LG2026
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
SangLyul Cho, Langqing Cui, Sehoon Kim +2
Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur…
cs.LG2024★ 4 cited
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
Yeonhong Park, Jake Hyun, SangLyul Cho +2
Recently, considerable efforts have been directed towards compressing Large Language Models (LLMs), which showcase groundbreaking capabilities across diverse applications but entai…