3 papers
cs.DC2026
NIXT: A NCCL Inspector Exporter Tool for Observability of Collective Communication in Large Model Training
Ziyang Jia, Sirshak Das, Jason Sewall +3
As machine learning workloads scale, it is increasingly important to gain more observability into the performance of collective communication to easily identify performance vari- a…
cs.DC2026
Energy-Efficient Multimodal Inference Serving with Tri-serve
Ziyang Jia, Sara Rashidi Golrouye, Laxmi Bhuyan +5
Multimodal model inference creates substantial energy demand with growing performance requirements. Within GPUs, power is autonomously managed by an on-board power management unit…
cs.LG2026
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
Lei Gao, Chaoyi Jiang, Hossein Entezari Zarch +3
Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode ph…