3 papers
cs.DC2026
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
Hongyi Jin, Bohan Hou, Guanjie Wang +18
Modern GPU workloads, especially large language model (LLM) inference, suffer from kernel launch overheads and coarse synchronization that limit inter-kernel parallelism. Recent me…
cs.DC2026
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
Bohan Hou, Hongyi Jin, Guanjie Wang +7
Scaling modern deep learning workloads demands coordinated placement of data and compute across device meshes, memory hierarchies, and heterogeneous accelerators. We present Axe La…
math.AP2024
On a Divergence Penalized Landau-de Gennes Model
Lia Bronsard, Jinqi Chen, Léa Mazzouza +4
We give a brief introduction to a divergence penalized Landau-de Gennes functional as a toy model for the study of nematic liquid crystal with colloid inclusion, in the case of une…