6 papers
μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAP
Shixin Ji, Jinming Zhuang, Zhuoping Yang +3
Heterogeneous reconfigurable platforms with tensor cores, such as AMD ACAP, are increasingly adopted for deep neural network (DNN) inference due to their high throughput and flexib…
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
Xingzhen Chen, Zhuoping Yang, Jinming Zhuang +5
As deep neural networks develop significantly more diverse and complex, achieving high performance and efficiency on complicated DNN models faces pressing challenges. Modern DNN wo…
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
Xingzhen Chen, Jinming Zhuang, Zhuoping Yang +5
With the development of deep neural network (DNN) enabled applications, achieving high hardware resource efficiency on diverse workloads is non-trivial in heterogeneous computing p…
PHAROS: Pipelined Heterogeneous Accelerators for Real-time Safety-critical Systems With Deadline Compliance
Shixin Ji, Jinming Zhuang, Sarah Schultz +6
Spatially partitioned heterogeneous accelerators (HAs) are increasingly adopted in embedded systems for their performance and flexibility. Yet most existing HA design frameworks op…
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
Erwei Wang, Samuel Bayliss, Andra Bisca +19
General-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures…
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
Zhuoping Yang, Jinming Zhuang, Xingzhen Chen +2
GPUs are critical for compute-intensive applications, yet emerging workloads such as recommender systems, graph analytics, and data analytics often exceed GPU memory capacity. Exis…