5 papers
Clean Me If You Can: A Large Collection of Real-World Addresses for Data Cleaning Benchmarking
Fatemeh Ahmadi, Tobias Bernhard, Mohamed Abdelmaksoud +3
There has been extensive research on automating and scaling data cleaning, i.e., the detection and correction of erroneous values in tabular data. Yet, existing approaches often pe…
Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement
Nils Strassenburg, Boris Glavic, Tilmann Rabl
Businesses increasingly rely on large language models (LLMs) to automate simple repetitive tasks instead of developing custom machine learning models. LLMs require few, if any, tra…
GenIE - Simulator-Driven Iterative Data Exploration for Scientific Discovery
Ashwin Gerard Colaco, Martin Boissier, Sriram Rao +3
Physics-based simulators play a critical role in scientific discovery and risk assessment, enabling what-if analyses for events like wildfires and hurricanes. Today, databases trea…
Skyrise: Exploiting Serverless Cloud Infrastructure for Elastic Data Processing
Thomas Bodner, Daniel Ritter, Martin Boissier +1
Serverless computing offers elasticity unmatched by conventional server-based cloud infrastructure. Although modern data processing systems embrace serverless storage, such as Amaz…
An Empirical Evaluation of Serverless Cloud Infrastructure for Large-Scale Data Processing
Thomas Bodner, Theo Radig, David Justen +2
Data processing systems are increasingly deployed in the cloud. While monolithic systems run fully on virtual servers, recent systems embrace cloud infrastructure and utilize the d…