3 papers
cs.AI2026
Parallel Context Compaction for Long-Horizon LLM Agent Serving
Musa Cim, Burak Topcu, Chita Das +1
Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based summarization keeps the conver…
cs.PF2025
GPU Cluster Scheduling for Network-Sensitive Deep Learning
Aakash Sharma, Vivek M. Bhasi, Sonali Singh +3
We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the…
cs.AR2024
Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
Rishabh Jain, Vivek M. Bhasi, Adwait Jog +3
Personalized recommendation is a ubiquitous application on the internet, with many industries and hyperscalers extensively leveraging Deep Learning Recommendation Models (DLRMs) fo…