2 citations · 2 across the 6 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
Chen Zhang, Yan Ding, Haotian Wang +3
During the deployment of Large Language Models (LLMs), the autoregressive decoding phase on heterogeneous NPU platforms (e.g., Ascend 910B) faces severe memory-bound challenges. Th…
cs.DC2022
cuFasterTucker: A Stochastic Optimization Strategy for Parallel Sparse FastTucker Decomposition on GPU Platform
Zixuan Li
Currently, the size of scientific data is growing at an unprecedented rate. Data in the form of tensors exhibit high-order, high-dimensional, and highly sparse features. Although t…