From the 1 of 4 linked papers with an AI index.
4 papers
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems
Jiayi Li, Di Wu, Qingxu Li +10
The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communicati…
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing
Daijing Shi, Hongxiao Zhao, Yihan Fu +7
Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequen…
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
Jiaqi Yang, Jiayi Li, Yihan Fu +5
The paper introduces DOPS, a framework that dynamically schedules LLM operators and chooses efficient weight layouts to improve inference latency on heterogeneous systems with NPUs…
RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression
Yuchao Qin, Anjunyi Fan, Bonan Yan
Data centers handle vast volumes of data that require efficient lossless compression, yet emerging probabilistic models based methods are often computationally slow. To address thi…