4 papers
C2P-Cache: Scalable GPU L1 Cache Sharing via Concurrent Candidate Pruning
Hanqing Li, Lizhou Wu, Tiejun Li +6
Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested l…
NeuroPDE+: A Scalable Neuromorphic PDE Accelerator Based on Spintronic and Ferroelectric Devices
Siqing Fu, Lizhou Wu, Tiejun Li +8
The pursuit of high-performance PDE solvers rests on three fundamental challenges: (i) the curse of dimensionality in kinetic and financial equations, (ii) the poor extrapolation o…
Spin-NeuroMem: A Low-Power Neuromorphic Associative Memory Design Based on Spintronic Devices
Siqing Fu, Lizhou Wu, Tiejun Li +3
Biologically-inspired computing models have made significant progress in recent years, but the conventional von Neumann architecture is inefficient for the large-scale matrix opera…
NeuroPDE: A Neuromorphic PDE Solver Based on Spintronic and Ferroelectric Devices
Siqing Fu, Lizhou Wu, Tiejun Li +5
In recent years, new methods for solving partial differential equations (PDEs) such as Monte Carlo random walk methods have gained considerable attention. However, due to the lack…