3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2026
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
Walaa Amer, Uday das, Fadi Kurdahi
Self-speculative decoding is an inference technique for large language models designed to speed up generation without sacrificing output quality. It combines fast, approximate deco…
cs.CL2024★ 3 cited
Evaluating the Performance of LLMs on Technical Language Processing tasks
Andrew Kernycky, David Coleman, Christopher Spence +1
In this paper we present the results of an evaluation study of the perfor-mance of LLMs on Technical Language Processing tasks. Humans are often confronted with tasks in which they…