2 papers
cs.AR2026
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
George Karfakis, Samyak Chakrabarty, Vinod Kurian Jacob +4
The proliferation of Transformer models is often constrained by the significant computational and memory bandwidth demands of deployment. To address this, we present MXFormer, a no…
cs.PF2025
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
George Karfakis, Faraz Tahmasebi, Binglu Chen +5
RAPID-LLM is a unified performance modeling framework for distributed large language model (LLM) training and inference on GPU clusters, without relying on deployment-specific trac…