1 paper · 1 filter
Haifeng Wu, Srinivasan Manoharan, Fangbo Tu +2
We present RLM-Cascade, a proxy-layer system that applies speculative decoding at the response level to reduce LLM API costs without requiring model architecture access or a shared…