3 papers
cs.CL2026
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
Vasu Shyam, Anna Golubeva, Quentin Anthony
We present tensor and sequence parallelism (TSP), a parallel execution strategy that folds tensor parallelism and sequence parallelism onto a single device axis. In conventional mu…
cs.CL2025
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Quentin Anthony, Yury Tokpanov, Skyler Szot +18
We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance…
cs.LG2024
The Zamba2 Suite: Technical Report
Paolo Glorioso, Quentin Anthony, Yury Tokpanov +5
In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance again…