works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.DC2026

C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems

Jiayi Li, Di Wu, Qingxu Li +10

The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communicati…

cs.AR2026

MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing

Daijing Shi, Hongxiao Zhao, Yihan Fu +7

Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequen…

cs.AR2026

Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling

Jiaqi Yang, Jiayi Li, Yihan Fu +5

The paper introduces DOPS, a framework that dynamically schedules LLM operators and chooses efficient weight layouts to improve inference latency on heterogeneous systems with NPUs…

cs.AR2025

RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression

Yuchao Qin, Anjunyi Fan, Bonan Yan

Data centers handle vast volumes of data that require efficient lossless compression, yet emerging probabilistic models based methods are often computationally slow. To address thi…

cs.AR2024

Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators

Ruibao Wang, Bonan Yan

Processing-in-memory (PIM) is a promising choice for accelerating deep neural networks (DNNs) featuring high efficiency and low power. However, the rapid upscaling of neural networ…