9 papers
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Stateful serverless (Function-as-a-Service) environments, whose workers host state servers, are increasingly used for stream processing. A stream application is a pipeline of opera…
AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows
Riya Samanta, Bidyut Saha, Soumya Kanti Ghosh +1
Modern LLM-based applications increasingly require multi- stage execution, persistent intermediate state, retry seman- tics, and auditable usage accounting. However, many LLM appli…
How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits
Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya
Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling…
FedCARE: A Multi-Objective Personalised Federated Learning Framework for Smart Healthcare
Rojalini Tripathy, Padmalochan Bera, Shreya Ghosh +1
Federated Learning (FL) enables collaborative model training across distributed healthcare institutions without centralising sensitive patient data. However, real-world healthcare…
Trust-Aware Topology Learning for Dynamic Decentralized Federated Learning under Adversaries
Shubham Vaishnav, Murtaza Rangwala, Ali Beikmohammadi +2
In dynamic mobile decentralized federated learning (DFL), adversaries can poison both model updates and the topology information devices use to choose collaborators. We present DMT…
Preserving Admission Responsibility in Multi-Tenant Large Language Model Prefix Caches
Zhiyu Wang, Rajkumar Buyya
Shared prefix caching turns Graphics Processing Unit (GPU) memory into persistent state shared across Large Language Model (LLM) tenants. A group that materializes new Key-Value (K…