activity
20172020
most citedExploring Erasure Coding Techniques for High Availability of Intermediate Data

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.DC20201 cited

Exploring Erasure Coding Techniques for High Availability of Intermediate Data

Zhe Zhang, Brian Bockelman, Derek Weitzel +1

Scientific computing workflows generate enormous distributed data that is short-lived, yet critical for job completion time. This class of data is called intermediate data. A commo…

cs.DC2020

Trua: Efficient Task Replication for Flexible User-defined Availability in Scientific Grids

Zhe Zhang, Brian Bockelman, Derek Weitzel +3

Failure is inevitable in scientific computing. As scientific applications and facilities increase their scales over the last decades, finding the root cause of a failure can be ver…

cs.OH2020

ROOT I/O compression improvements for HEP analysis

Oksana Shadura, Brian Paul Bockelman, Philippe Canal +2

We overview recent changes in the ROOT I/O system, increasing performance and enhancing it and improving its interaction with other data analysis ecosystems. Both the newly introdu…

cs.DC2019

Speeding HEP Analysis with ROOT Bulk I/O

Brian Bockelman, Zhe Zhang, Oksana Shadura

Distinct HEP workflows have distinct I/O needs; while ROOT I/O excels at serializing complex C++ objects common to reconstruction, analysis workflows typically have simpler objects…

cs.PL2017

Fast Access to Columnar, Hierarchically Nested Data via Code Transformation

Jim Pivarski, Peter Elmer, Brian Bockelman +1

Big Data query systems represent data in a columnar format for fast, selective access, and in some cases (e.g. Apache Drill), perform calculations directly on the columnar data wit…