1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.AI2025
Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments
Xinkai Zou, Xuan Jiang, Ruikai Huang +8
Anomaly detection in cloud environments remains both critical and challenging. Existing context-level benchmarks typically focus on either metrics or logs and often lack reliable a…
cs.DC2024★ 1 cited
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models
Jialiang Cheng, Ning Gao, Yun Yue +3
Distributed training methods are crucial for large language models (LLMs). However, existing distributed training methods often suffer from communication bottlenecks, stragglers, a…
cs.DB2024
Couler: Unified Machine Learning Workflow Optimization in Cloud
Xiaoda Wang, Yuan Tang, Tengda Guo +6
Machine Learning (ML) has become ubiquitous, fueling data-driven applications across various organizations. Contrary to the traditional perception of ML in research, ML workflows c…