4 papers
A Real-Time Digital Twin for Adaptive Scheduling
Yihe Zhang, Yash Kurkure, Yiheng Tao +2
High-performance computing (HPC) workloads are becoming increasingly diverse, exhibiting wide variability in job characteristics, yet cluster scheduling has long relied on static,…
Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
Yiheng Tao, Yihe Zhang, Matthew Dearing +4
Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with t…
Leveraging LLMs to Automate Energy-Aware Refactoring of Parallel Scientific Codes
Matthew T. Dearing, Yiheng Tao, Xingfu Wu +2
Large language models (LLMs) are increasingly used for generating parallel scientific codes, with a primary focus on generating functionally correct code. Recent work has focused o…
LASSI: An LLM-based Automated Self-Correcting Pipeline for Translating Parallel Scientific Codes
Matthew T. Dearing, Yiheng Tao, Xingfu Wu +2
This paper addresses the problem of providing a novel approach to sourcing significant training data for LLMs focused on science and engineering. In particular, a crucial challenge…