1 citations · 1 across the 3 of their papers we have counts for
4 papers
Improving scalability and reliability of MPI-agnostic transparent checkpointing for production workloads at NERSC
Prashant Singh Chouhan, Harsh Khetawat, Neil Resnik +5
Checkpoint/restart (C/R) provides fault-tolerant computing capability, enables long running applications, and provides scheduling flexibility for computing centers to support diver…
Checkpointing SPAdes for Metagenome Assembly: Transparency versus Performance in Production
Twinkle Jain, Jie Wang
The SPAdes assembler for metagenome assembly is a long-running application commonly used at the NERSC supercomputing site. However, NERSC, like many other sites, has a 48-hour limi…
CRAC: Checkpoint-Restart Architecture for CUDA with Streams and UVM
Twinkle Jain, Gene Cooperman
The share of the top 500 supercomputers with NVIDIA GPUs is now over 25% and continues to grow. While fault tolerance is a critical issue for supercomputing, there does not current…
Data Comets: Designing a Visualization Tool for Analyzing Autonomous Aerial Vehicle Logs with Grounded Evaluation
David Saffo, Aristotelis Leventidis, Twinkle Jain +2
Autonomous unmanned aerial vehicles are complex systems of hardware, software, and human input. Understanding this complexity is key to their development and operation. Information…