1 paper · 1 filter
Ali Noshad, Zishan Zheng, Yinjun Wu
To reduce LLM costs and latency, semantic caching systems must accurately identify when a new prompt matches a cached one. Current methods often rely on simplistic similarity measu…