1 paper
Rajveer Bachkaniwala, Chengqi Luo, Richard So +2
Context retrieval systems for LLM inference face a critical challenge: high retrieval latency creates a fundamental tension between waiting for complete context (poor time-to-first…