2 papers
cs.AR2026
μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAP
Shixin Ji, Jinming Zhuang, Zhuoping Yang +3
Heterogeneous reconfigurable platforms with tensor cores, such as AMD ACAP, are increasingly adopted for deep neural network (DNN) inference due to their high throughput and flexib…
cs.DC2026
LiveR: Fine-Grained Elasticity via Live Reconfiguration for Model Training
Haoyuan Liu, Kairui Zhou, Shuyao Qi +4
To reduce user costs and maximize cluster utilization, large model training increasingly leverages volatile but inexpensive GPU capacity, such as spot instances and reclaimable res…