126 citations · 146 across the 9 of their papers we have counts for
5 papers · 1 filter
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
Bole Ma, Jan Eitzinger, Harald Köstler +1
Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit is now a small, reusable chun…
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
Bole Ma, Jan Eitzinger, Harald Koestler +1
AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement,…
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
Philipp Suffa, Markus Holzer, Harald Köstler +1
We implement and analyse a sparse / indirect-addressing data structure for the Lattice Boltzmann Method to support efficient compute kernels for fluid dynamics problems with a high…
waLBerla: A block-structured high-performance framework for multiphysics simulations
Martin Bauer, Sebastian Eibl, Christian Godenschwager +8
Programming current supercomputers efficiently is a challenging task. Multiple levels of parallelism on the core, on the compute node, and between nodes need to be exploited to mak…
Massively Parallel Phase-Field Simulations for Ternary Eutectic Directional Solidification
Martin Bauer, Johannes Hötzer, Philipp Steinmetz +7
Microstructures forming during ternary eutectic directional solidification processes have significant influence on the macroscopic mechanical properties of metal alloys. For a real…