collaborators

5 papers

cs.DC2026

Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention

Dingyu Yang, Fanyong Kong, Jie Dai +5

Modern cloud servers routinely co-locate multiple latency-sensitive microservice instances to improve resource efficiency. However, the diversity of microservice behaviors, coupled…

cs.LG2025

EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training

Qingao Yi, Jiaang Duan, Hanwen Hu +10

Training large language models (LLMs) poses significant challenges regarding computational resources and memory capacity. Although distributed training techniques help mitigate the…

cs.DC2025

GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management

Jiaang Duan, Shenglin Xu, Shiyou Qian +15

The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cl…

cs.PF2025

Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices

Jiaqi Sun, Dingyu Yang, Shiyou Qian +2

To handle the high volume of requests, large-scale services are comprised of thousands of instances deployed in clouds. These services utilize diverse programming languages and are…

cs.CL2025

LKD-KGC: Domain-Specific KG Construction via LLM-driven Knowledge Dependency Parsing

Jiaqi Sun, Shiyou Qian, Zhangchi Han +5

Knowledge Graphs (KGs) structure real-world entities and their relationships into triples, enhancing machine reasoning for various tasks. While domain-specific KGs offer substantia…