collaborators

5 papers

cs.CR2026

Hazel: Secure and Efficient Disaggregated Storage

Marcin Chrapek, Meni Orenbach, Ahmad Atamli +4

Disaggregated storage with NVMe-over-Fabrics (NVMe-oF) has emerged as the standard solution in modern supercomputers and data center clusters, achieving superior performance, resou…

cs.NI2026

REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure Mitigation

Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +7

Next-generation datacenters require highly efficient network load balancing to manage the growing scale of artificial intelligence (AI) training and general datacenter traffic. How…

cs.NI2026

In-Network Collective Operations: Game Changer or Challenge for AI Workloads?

Torsten Hoefler, Mikhail Khalilov, Josiah Clark +9

This paper summarizes the opportunities of in-network collective operations (INC) for accelerated collective operations in AI workloads. We provide sufficient detail to make this i…

cs.PF2025

EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC

Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6

Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…

cs.NI2025

SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication

Mikhail Khalilov, Siyuan Shen, Marcin Chrapek +16

RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-…