Showing cs.PFShow all
2 papers · 1 filter
cs.PF2025★ 1 cited
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6
Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…
cs.PF2025
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
Marcin Chrapek, Marcin Copik, Etienne Mettaz +1
Large Language Models (LLMs) are increasingly deployed on converged Cloud and High-Performance Computing (HPC) infrastructure. However, as LLMs handle confidential inputs and are f…