collaborators

8 papers

cs.DB2026

Benchmarking Filtered Approximate Nearest Neighbor Search Algorithms on Transformer-based Embedding Vectors

Patrick Iff, Paul Bruegger, Marcin Chrapek +3

Advances in embedding models for text, image, audio, and video drive progress across multiple domains, including retrieval-augmented generation, recommendation systems, and others.…

cs.CL2026

Multi-Head RAG: Solving Multi-Aspect Problems with LLMs

Maciej Besta, Ales Kubicek, Robert Gerstenberger +13

Retrieval-Augmented Generation (RAG) improves Large Language Models (LLMs) by retrieving supporting documents into the prompt, but existing methods do not explicitly target queries…

cs.CR2026

Hazel: Secure and Efficient Disaggregated Storage

Marcin Chrapek, Meni Orenbach, Ahmad Atamli +4

Disaggregated storage with NVMe-over-Fabrics (NVMe-oF) has emerged as the standard solution in modern supercomputers and data center clusters, achieving superior performance, resou…

cs.PF2025

EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC

Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6

Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…

cs.PF2025

PerfDojo: Automated ML Library Generation for Heterogeneous Architectures

Andrei Ivanov, Siyuan Shen, Gioele Gottardo +5

The increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a signifi…

cs.PF2025

Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs

Marcin Chrapek, Marcin Copik, Etienne Mettaz +1

Large Language Models (LLMs) are increasingly deployed on converged Cloud and High-Performance Computing (HPC) infrastructure. However, as LLMs handle confidential inputs and are f…