10 papers · 1 filter
Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions
Niloofar Gholipour, Marcos Assuncao, Gursimran Singh +8
Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training…
Agentic Autoscaling through Worker-Pool Orchestration for LLM-driven Text Classification in Cloud Computing Environments
Bablu Kumar, Anshul Verma, Rajkumar Buyya
The growing adoption of large language model (LLM)-based systems for large-scale text processing has created a critical need for dynamic autoscaling to manage high-latency, bursty,…
Stability-Aware Proactive Autoscaling Using a Double Deep Q-Network in Cloud Computing Environments
Bablu Kumar, Anshul Verma, Rajkumar Buyya
Dynamic workloads and latency-sensitive applications require efficient autoscaling in cloud computing environments. However, most existing approaches rely on reactive mechanisms ba…
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Stateful serverless (Function-as-a-Service) environments, whose workers host state servers, are increasingly used for stream processing. A stream application is a pipeline of opera…
AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows
Riya Samanta, Bidyut Saha, Soumya Kanti Ghosh +1
Modern LLM-based applications increasingly require multi- stage execution, persistent intermediate state, retry seman- tics, and auditable usage accounting. However, many LLM appli…
How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits
Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya
Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling…