Showing cs.OSShow all
2 papers · 1 filter
cs.OS2026
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
Yifei Wang, Hancheng Ye, Yechen Xu +8
Large language models (LLMs) are increasingly deployed as the execution core of autonomous agents rather than as standalone text generators. Agentic workloads induce a temporal shi…
cs.OS2026
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
Yechen Xu, Yifei Wang, Nathanael Ren +2
Consumer machines are increasingly running large ML workloads such as large language models (LLMs), text-to-image generation, and interactive image editing. Unlike datacenter GPUs,…