most citedEDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.PF20251 cited

EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC

Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6

Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…

cs.PF2025

PerfDojo: Automated ML Library Generation for Heterogeneous Architectures

Andrei Ivanov, Siyuan Shen, Gioele Gottardo +5

The increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a signifi…

cs.PF2025

Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs

Marcin Chrapek, Marcin Copik, Etienne Mettaz +1

Large Language Models (LLMs) are increasingly deployed on converged Cloud and High-Performance Computing (HPC) infrastructure. However, as LLMs handle confidential inputs and are f…

cs.AI20251 cited

Psychologically Enhanced AI Agents

Maciej Besta, Shriram Chandran, Robert Gerstenberger +9

We introduce MBTI-in-Thoughts, a framework for enhancing the effectiveness of Large Language Model (LLM) agents through psychologically grounded personality conditioning. Drawing o…

cs.NI2025

SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication

Mikhail Khalilov, Siyuan Shen, Marcin Chrapek +16

RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-…