3 papers
cs.LG2026
FlashAttention for Scalable Vector Architectures
Sonia Rani Gupta, Nikela Papadopoulou, Miquel Pericàs
Inference with transformer models on CPUs is increasingly important, especially for Small Language Models (SLMs), where vector architectures are emerging as a promising execution s…
cs.DC2026
DEFT: Joint Task Placement and DVFS for Energy-Efficient Multi-GPU Runtimes
Jing Chen, Miquel Pericas
Energy efficiency has become a first-order concern in modern high-performance computing systems, as it directly determines achievable throughput under fixed power budgets. Although…
cs.DC2026
Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads
Minyu Cui, Miquel Pericas
The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems. As model sizes and computatio…