1 paper
Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev +2
Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only m…