3 papers
cs.DC2025
Flora: Efficient Cloud Resource Selection for Big Data Processing via Job Classification
Jonathan Will, Lauritz Thamsen, Jonathan Bader +1
Distributed dataflow systems like Spark and Flink enable data-parallel processing of large datasets on clusters of cloud resources. Yet, selecting appropriate computational resourc…
cs.DC2025
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
Jonathan Bader, Fabian Skalski, Fabian Lehmann +4
As the amount of available data continues to grow in fields as diverse as bioinformatics, physics, and remote sensing, the importance of scientific workflows in the design and impl…
cs.DC2025
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
Jonathan Will, Nico Treide, Lauritz Thamsen +1
Distributed dataflow systems like Spark and Flink enable data-parallel processing of large datasets on clusters. Yet, selecting appropriate computational resources for dataflow job…