2 papers
cs.AR2026
C2P-Cache: Scalable GPU L1 Cache Sharing via Concurrent Candidate Pruning
Hanqing Li, Lizhou Wu, Tiejun Li +6
Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested l…
cs.AR2026
NeuroPDE+: A Scalable Neuromorphic PDE Accelerator Based on Spintronic and Ferroelectric Devices
Siqing Fu, Lizhou Wu, Tiejun Li +8
The pursuit of high-performance PDE solvers rests on three fundamental challenges: (i) the curse of dimensionality in kinetic and financial equations, (ii) the poor extrapolation o…