papers

Publications (13)

cs.DB2019

Cloudy with high chance of DBMS: A 10-year prediction for Enterprise-Grade ML

Ashvin Agrawal, Rony Chatterjee, Carlo Curino +19

Machine learning (ML) has proven itself in high-value web applications such as search ranking and is emerging as a powerful tool in a much broader range of enterprise scenarios inc…

cs.AI2026

GraphMind: From Operational Traces to Self-Evolving Workflow Automation

Yiwen Zhu, Joyce Cahoon, Anna Pavlenko +13

Complex operational workflows coordinating personnel, tools, and information are central to system operations, yet end-to-end automation remains challenging due to extensive human…

cs.DB2024

Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service

Deepak Ravikumar, Alex Yeo, Yiwen Zhu +11

The proliferation of big data and analytic workloads has driven the need for cloud compute and cluster-based job processing. With Apache Spark, users can process terabytes of data…

cs.DB2024

Lorentz: Learned SKU Recommendation Using Profile Data

Nicholas Glaze, Tria McNeely, Yiwen Zhu +4

Cloud operators have expanded their service offerings, known as Stock Keeping Units (SKUs), to accommodate diverse demands, resulting in increased complexity for customers to selec…

cs.DB2021

KEA: Tuning an Exabyte-Scale Data Infrastructure

Yiwen Zhu, Subru Krishnan, Konstantinos Karanasos +12

Microsoft's internal big-data infrastructure is one of the largest in the world -- with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infra…

cs.DB2022

Doppler: Automated SKU Recommendation in Migrating SQL Workloads to the Cloud

Joyce Cahoon, Wenjing Wang, Yiwen Zhu +8

Selecting the optimal cloud target to migrate SQL estates from on-premises to the cloud remains a challenge. Current solutions are not only time-consuming and error-prone, requirin…

cs.LG2019

Griffon: Reasoning about Job Anomalies with Unlabeled Data in Cloud-based Platforms

Liqun Shao, Yiwen Zhu, Abhiram Eswaran +9

Microsoft's internal big data analytics platform is comprised of hundreds of thousands of machines, serving over half a million jobs daily, from thousands of users. The majority of…

cs.LG2020

Vamsa: Automated Provenance Tracking in Data Science Scripts

Mohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas +5

There has recently been a lot of ongoing research in the areas of fairness, bias and explainability of machine learning (ML) models due to the self-evident or regulatory requiremen…

cs.LG2019

Data Science through the looking glass and what we found there

Fotis Psallidas, Yiwen Zhu, Bojan Karlas +8

The recent success of machine learning (ML) has led to an explosive growth both in terms of new systems and algorithms built in industry and academia, and new applications built by…

cs.SE2026

ENCO: Life-Cycle Management of Enterprise-Grade Copilots

Yiwen Zhu, Mathieu Demarne, Kai Deng +12

Software engineers frequently grapple with the challenge of accessing disparate documentation and telemetry data, including TroubleShooting Guides (TSGs), incident reports, code re…

cs.IR2025

FLAIR: Feedback Learning for Adaptive Information Retrieval

William Zhang, Yiwen Zhu, Yunlei Lu +7

Recent advances in Large Language Models (LLMs) have driven the adoption of copilots in complex technical scenarios, underscoring the growing need for specialized information retri…

cs.DB2019

Extending Relational Query Processing with ML Inference

Konstantinos Karanasos, Matteo Interlandi, Doris Xin +10

The broadening adoption of machine learning in the enterprise is increasing the pressure for strict governance and cost-effective performance, in particular for the common and cons…

cs.DC2024

Towards Building Autonomous Data Services on Azure

Yiwen Zhu, Yuanyuan Tian, Joyce Cahoon +35

Modern cloud has turned data services into easily accessible commodities. With just a few clicks, users are now able to access a catalog of data processing systems for a wide range…