dynamic operator scheduling 1heterogeneous platforms 1large language model inference 1neural processing units 1processing-in-memory 1weight layout optimization 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.DC2026
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems
Jiayi Li, Di Wu, Qingxu Li +10
The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communicati…
cs.AR2026
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
Jiaqi Yang, Jiayi Li, Yihan Fu +5
The paper introduces DOPS, a framework that dynamically schedules LLM operators and chooses efficient weight layouts to improve inference latency on heterogeneous systems with NPUs…
cs.LG2025
Language-Enhanced Representation Learning for Single-Cell Transcriptomics
Yaorui Shi, Jiaqi Yang, Changhao Nai +5
Single-cell RNA sequencing (scRNA-seq) offers detailed insights into cellular heterogeneity. Recent advancements leverage single-cell large language models (scLLMs) for effective r…