3 papers
cs.DB2025
Vortex: Hosting ML Inference and Knowledge Retrieval Services With Tight Latency and Throughput Requirements
Yuting Yang, Tiancheng Yuan, Jamal Hashim +6
There is growing interest in deploying ML inference and knowledge retrieval as services that could support both interactive queries by end users and more demanding request flows th…
cs.DC2024
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
Yuting Yang, Andrea Merlina, Weijia Song +3
We consider ML query processing in distributed systems where GPU-enabled workers coordinate to execute complex queries: a computing style often seen in applications that interact w…
cs.OS2023
Cascade: A Platform for Delay-Sensitive Edge Intelligence
Weijia Song, Thiago Garrett, Yuting Yang +6
Interactive intelligent computing applications are increasingly prevalent, creating a need for AI/ML platforms optimized to reduce per-event latency while maintaining high throughp…