3 citations · 4 across the 3 of their papers we have counts for
6 papers
Algorithms for a Topology-aware Massively Parallel Computation Model
Xiao Hu, Paraschos Koutris, Spyros Blanas
Most of the prior work in massively parallel data processing assumes homogeneity, i.e., every computing unit has the same computational capability, and can communicate with every o…
Chasing Similarity: Distribution-aware Aggregation Scheduling (Extended Version)
Feilong Liu, Ario Salmasi, Spyros Blanas +1
Parallel aggregation is a ubiquitous operation in data analytics that is expressed as GROUP BY in SQL, reduce in Hadoop, or segment in TensorFlow. Parallel aggregation starts with…
To Ship or Not to (Function) Ship (Extended version)
Feilong Liu, Niranjan Kamat, Spyros Blanas +1
Sampling is often used to reduce query latency for interactive big data analytics. The established parallel data processing paradigm relies on function shipping, where a coordinato…
Approximate Distributed Joins in Apache Spark
Do Le Quoc, Istemi Ekin Akkus, Pramod Bhatotia +4
The join operation is a fundamental building block of parallel data processing. Unfortunately, it is very resource-intensive to compute an equi-join across massive datasets. The ap…
ArrayBridge: Interweaving declarative array processing with high-performance computing
Haoyuan Xing, Sofoklis Floratos, Spyros Blanas +4
Scientists are increasingly turning to datacenter-scale computers to produce and analyze massive arrays. Despite decades of database research that extols the virtues of declarative…
Towards Exascale Scientific Metadata Management
Spyros Blanas, Surendra Byna
Advances in technology and computing hardware are enabling scientists from all areas of science to produce massive amounts of data using large-scale simulations or observational fa…