Modyn: Data-Centric Machine Learning Pipeline Orchestration
arXiv:2312.06254 · doi:10.1145/3709705
Abstract
In real-world machine learning (ML) pipelines, datasets are continuously growing. Models must incorporate this new training data to improve generalization and adapt to potential distribution shifts. The cost of model retraining is proportional to how frequently the model is retrained and how much data it is trained on, which makes the naive approach of retraining from scratch each time impractical. We present Modyn, a data-centric end-to-end machine learning platform. Modyn's ML pipeline abstraction enables users to declaratively describe policies for continuously training a model on a growing dataset. Modyn pipelines allow users to apply data selection policies (to reduce the number of data points) and triggering policies (to reduce the number of trainings). Modyn executes and orchestrates these continuous ML training pipelines. The system is open-source and comes with an ecosystem of benchmark datasets, models, and tooling. We formally discuss how to measure the performance of ML pipelines by introducing the concept of composite models, enabling fair comparison of pipelines with different data selection and triggering policies. We empirically analyze how various data selection and triggering policies impact model accuracy, and also show that Modyn enables high throughput training with sample-level data selection.
final version published at SIGMOD'25; 30 pages
References in corpus (22)
- AutoML: A Survey of the State-of-the-Art
- Learning under Concept Drift: A Review
- A Survey of Human-in-the-loop for Machine Learning
- Challenges in Deploying Machine Learning: a Survey of Case Studies
- Exponentially Weighted Moving Average Charts for Detecting Concept Drift
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- The Pipeline for the Continuous Development of Artificial Intelligence Models -- Current State of Research and Practice
- Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training
- Accelerating Recommendation System Training by Leveraging Popular Choices
- Vamsa: Automated Provenance Tracking in Data Science Scripts
- PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models
- "We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning
- Continual Learning in Practice
- Online Learning for Recommendations at Grubhub
- Data Management For Training Large Language Models: A Survey
- Online Continual Learning Without the Storage Constraint
- Rethinking Streaming Machine Learning Evaluation
- Less is more: Selecting informative and diverse subsets with balancing constraints
- A Framework for Monitoring and Retraining Language Models in Real-World Applications
- Optimization of Rank Losses for Image Retrieval
- MGit: A Model Versioning and Management System
- Cost-Effective Retraining of Machine Learning Models