Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs
arXiv:2108.11359 · doi:10.1109/HPEC49654.2021.9622870
Abstract
Diverse workloads such as interactive supercomputing, big data analysis, and large-scale AI algorithm development, requires a high-performance scheduler. This paper presents a novel node-based scheduling approach for large scale simulations of short running jobs on MIT SuperCloud systems, that allows the resources to be fully utilized for both long running batch jobs while simultaneously providing fast launch and release of large-scale short running jobs. The node-based scheduling approach has demonstrated up to 100 times faster scheduler performance that other state-of-the-art systems.
IEEE HPEC 2021
References in corpus (7)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- The Lustre Storage Architecture
- Computing on Masked Data: a High Performance Method for Improving Big Data Veracity
- D4M 2.0 Schema: A General Purpose High Performance Schema for the Accumulo Database
- Achieving 100,000,000 database inserts per second using Accumulo and D4M
- LLMapReduce: Multi-Level Map-Reduce for High Performance Data Analysis
- Best of Both Worlds: High Performance Interactive and Batch Launching