2 papers
cs.DB2025
Vortex: Hosting ML Inference and Knowledge Retrieval Services With Tight Latency and Throughput Requirements
Yuting Yang, Tiancheng Yuan, Jamal Hashim +6
There is growing interest in deploying ML inference and knowledge retrieval as services that could support both interactive queries by end users and more demanding request flows th…
cs.AI2025
Diagnosing and Resolving Cloud Platform Instability with Multi-modal RAG LLMs
Yifan Wang, Kenneth P. Birman
Today's cloud-hosted applications and services are complex systems, and a performance or functional instability can have dozens or hundreds of potential root causes. Our hypothesis…