1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2025
Fast Inference via Hierarchical Speculative Decoding
Clara Mohri, Haim Kaplan, Tal Schuster +2
Transformer language models generate text autoregressively, making inference latency proportional to the number of tokens generated. Speculative decoding reduces this latency witho…
cs.CY2025★ 1 cited
Towards an AI-Augmented Textbook
LearnLM Team, Google, : +34
Textbooks are a cornerstone of education, but they have a fundamental limitation: they are a one-size-fits-all medium. Any new material or alternative representation requires arduo…
cs.CL2024
Teaching Models to Improve on Tape
Liat Bezalel, Eyal Orgad, Amir Globerson
Large Language Models (LLMs) often struggle when prompted to generate content under specific constraints. However, in such cases it is often easy to check whether these constraints…