1 paper
Tanay Chowdhury, Saeideh Shahrokh Esfahani
Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds…