2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.DC2026
Quasar: Quantized Self-Speculative Acceleration for Rapid Inference via Memory-Efficient Verification
Guang Huang, Zeyi Wen
Speculative Decoding (SD) has emerged as a premier technique for accelerating Large Language Model (LLM) inference by decoupling token generation into rapid drafting and parallel v…
cs.AI2025★ 2 cited
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…