3 papers
cs.CL2024
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
David Koeplinger, Darshan Gandhi, Pushkar Nandkar +9
Token generation speed is critical to power the next wave of AI inference applications. GPUs significantly underperform during token generation due to synchronization overheads at…
cs.CR2024
PVF:Understanding AI Vulnerability Against SDCs
Xun Jiao, Fred Lin, Harish D. Dixit +8
Reliability of AI systems is a fundamental concern for the successful deployment and widespread adoption of AI technologies. Unfortunately, the escalating complexity and heterogene…
cs.OH2023
Manuscript of a method for improving wear in intermittently computing file systems
Yeteng Liao, Han wang
For the first time, the repeated wear phenomenon of high-frequency power failure on the data block area in intermittent computing file system is found. A method to improve NVM wear…