2 papers
cs.CL2025
The Disparate Impacts of Speculative Decoding
Jameson Sandler, Ahmet Üstün, Marco Romanelli +2
The practice of speculative decoding, whereby inference is probabilistically supported by a smaller, cheaper, ``drafter'' model, has become a standard technique for systematically…
cs.LG2024
Low-rank finetuning for LLMs: A fairness perspective
Saswat Das, Marco Romanelli, Cuong Tran +3
Low-rank approximation techniques have become the de facto standard for fine-tuning Large Language Models (LLMs) due to their reduced computational and memory requirements. This pa…