2 papers
cs.LG2026
SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems
Muhammad Faizan Raza, Shuo, Yang +1
Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval…
cs.LG2026
Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
Muhammad Faizan Raza, Shuo, Yang +2
Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pat…