collaborators

5 papers

cs.DB2026

Clean Me If You Can: A Large Collection of Real-World Addresses for Data Cleaning Benchmarking

Fatemeh Ahmadi, Tobias Bernhard, Mohamed Abdelmaksoud +3

There has been extensive research on automating and scaling data cleaning, i.e., the detection and correction of erroneous values in tabular data. Yet, existing approaches often pe…

cs.DB2026

Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement

Nils Strassenburg, Boris Glavic, Tilmann Rabl

Businesses increasingly rely on large language models (LLMs) to automate simple repetitive tasks instead of developing custom machine learning models. LLMs require few, if any, tra…

cs.DB2025

GenIE - Simulator-Driven Iterative Data Exploration for Scientific Discovery

Ashwin Gerard Colaco, Martin Boissier, Sriram Rao +3

Physics-based simulators play a critical role in scientific discovery and risk assessment, enabling what-if analyses for events like wildfires and hurricanes. Today, databases trea…

cs.DB2025

Skyrise: Exploiting Serverless Cloud Infrastructure for Elastic Data Processing

Thomas Bodner, Daniel Ritter, Martin Boissier +1

Serverless computing offers elasticity unmatched by conventional server-based cloud infrastructure. Although modern data processing systems embrace serverless storage, such as Amaz…

cs.DB2025

An Empirical Evaluation of Serverless Cloud Infrastructure for Large-Scale Data Processing

Thomas Bodner, Theo Radig, David Justen +2

Data processing systems are increasingly deployed in the cloud. While monolithic systems run fully on virtual servers, recent systems embrace cloud infrastructure and utilize the d…