3 papers
cs.DC2026
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
Yujie Zhang, Huiying Lan, Ehsan Aghapour +5
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining tech…
cs.AR2026
A Data-Driven Dynamic Execution Orchestration Architecture
Zhenyu Bai, Pranav Dangi, Rohan Juneja +4
Domain-specific accelerators deliver exceptional performance on their target workloads through fabrication-time orchestrated datapaths. However, such specialized architectures ofte…
cs.DC2025
TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
Wei Li, Zhenyu Bai, Heru Wang +6
Spatial dataflow accelerators are a promising direction for next-generation computer systems because they can reduce the memory bottlenecks of traditional von Neumann machines such…