1 paper · 1 filter
Congchao Wang, Sean Augenstein, Keith Rush +5
Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications. To address this, cascades of LMs offer an effective…