1 paper · 1 filter
Rajveer Bachkaniwala, Chengqi Luo, Richard So +2
Context retrieval systems for LLM inference face a critical challenge: high retrieval latency creates a fundamental tension between waiting for complete context (poor time-to-first…