4 papers
Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
Yuchen Fan, Minghong Sun, Jikui Ma +19
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collec…
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
Zongle Huang, Hongyang Jia, Kaiwei Zou +1
Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resourc…
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
Xinyuan Lin, Chenlu Li, Zongle Huang +5
Larger model sizes and longer sequence lengths have empowered the Large Language Model (LLM) to achieve outstanding performance across various domains. However, this progress bring…
Hecaton: Training Large Language Models with Scalable Chiplet Systems
Zongle Huang, Shupei Fan, Chen Tang +3
Large Language Models (LLMs) have achieved remarkable success in various fields, but their training and finetuning require massive computation and memory, necessitating parallelism…