activity
20242026
collaborators
Showing cs.ARShow all

5 papers · 1 filter

cs.AR2026

μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAP

Shixin Ji, Jinming Zhuang, Zhuoping Yang +3

Heterogeneous reconfigurable platforms with tensor cores, such as AMD ACAP, are increasingly adopted for deep neural network (DNN) inference due to their high throughput and flexib…

cs.AR2026

DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration

Xingzhen Chen, Zhuoping Yang, Jinming Zhuang +5

As deep neural networks develop significantly more diverse and complex, achieving high performance and efficiency on complicated DNN models faces pressing challenges. Modern DNN wo…

cs.AR2026

To Overlay or to Customize? Revisiting Architectural Choices in Heterogeneous Systems

Xingzhen Chen, Shixin Ji, Zheng Dong +1

In this work, we present a systematic study of this trade-off from a deployment-centric perspective, focusing on an autonomous driving scenario. Instead of treating overlay and cus…

cs.AR2026

FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration

Xingzhen Chen, Jinming Zhuang, Zhuoping Yang +5

With the development of deep neural network (DNN) enabled applications, achieving high hardware resource efficiency on diverse workloads is non-trivial in heterogeneous computing p…

cs.AR2026

PHAROS: Pipelined Heterogeneous Accelerators for Real-time Safety-critical Systems With Deadline Compliance

Shixin Ji, Jinming Zhuang, Sarah Schultz +6

Spatially partitioned heterogeneous accelerators (HAs) are increasingly adopted in embedded systems for their performance and flexibility. Yet most existing HA design frameworks op…