10 citations · 10 across the 1 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2023★ 2 cited
Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
Jonathan Will, Lauritz Thamsen, Dominik Scheinert +1
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate…
cs.DC2022★ 10 cited
Reshi: Recommending Resources for Scientific Workflow Tasks on Heterogeneous Infrastructures
Jonathan Bader, Fabian Lehmann, Alexander Groth +5
Scientific workflows typically comprise a multitude of different processing steps which often are executed in parallel on different partitions of the input data. These executions,…