1 paper
Cornelius Kummer, Lena Jurkschat, Michael Färber +1
With the wide adoption of language models for IR -- and specifically RAG systems -- the latency of the underlying LLM becomes a crucial bottleneck, since the long contexts of retri…