5 papers
ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization
Weijia Song, Jiashu Yue, Zhe Pang
How should multi-agent systems be designed, and can that design knowledge be captured in a form that is inspectable, revisable, and transferable? We introduce ABSTRAL, a framework…
Vortex: Hosting ML Inference and Knowledge Retrieval Services With Tight Latency and Throughput Requirements
Yuting Yang, Tiancheng Yuan, Jamal Hashim +6
There is growing interest in deploying ML inference and knowledge retrieval as services that could support both interactive queries by end users and more demanding request flows th…
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
Yuting Yang, Andrea Merlina, Weijia Song +3
We consider ML query processing in distributed systems where GPU-enabled workers coordinate to execute complex queries: a computing style often seen in applications that interact w…
Keep Your Friends Close: Leveraging Affinity Groups to Accelerate AI Inference Workflows
Thiago Garrett, Weijia Song, Roman Vitenberg +1
AI inference workflows are typically structured as a pipeline or graph of AI programs triggered by events. As events occur, the AIs perform inference or classification tasks under…
Cascade: A Platform for Delay-Sensitive Edge Intelligence
Weijia Song, Thiago Garrett, Yuting Yang +6
Interactive intelligent computing applications are increasingly prevalent, creating a need for AI/ML platforms optimized to reduce per-event latency while maintaining high throughp…