activity
20192026
most citedMimose: An Input-Aware Checkpointing Planner for Efficient Training on GPU

2 citations · 4 across the 10 of their papers we have counts for

collaborators

14 papers

cs.DC2026

DB-SpMSpV: Dual-View Blocked Sparse Matrix-Sparse Vector Multiplication for Dynamic GPU Workloads

Xing Cong, Chenhao Xie, Rui Wang +3

Sparse Matrix-Sparse Vector Multiplication (SpMSpV) is a core primitive in graph traversal, sparse linear algebra, and sparse model inference. Its input vector is often dynamically…

cs.CL2026

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

Yuchen Wang, Zhongzhi Luan

Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing,…

cs.SE2026

Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning

Jiaxing Qi, Zhongzhi Luan, Hongyu Zhang +5

Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems. However, recovery-oriented u…

cs.DC2026

RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms

Yao Lu, Shiqing Ma, Zhongzhi Luan +5

Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typi…

cs.DC2025

PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization

Kelun Lei, Hailong Yang, Huaitao Zhang +5

Designing high-performance kernels requires expert-level tuning and a deep understanding of hardware characteristics. Recent advances in large language models (LLMs) have enabled a…

cs.DC2025

LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures

Kelun Lei, Hailong Yang, Kaige Zhang +8

Sparse matrix-dense matrix multiplication (SpMM) is a critical kernel in both scientific computing and emerging graph learning workloads. The recent Armv9 architecture introduces S…