1 paper
Ali Noshad, Zishan Zheng, Yinjun Wu
To reduce LLM costs and latency, semantic caching systems must accurately identify when a new prompt matches a cached one. Current methods often rely on simplistic similarity measu…