40 citations · 123 across the 10 of their papers we have counts for
9 papers · 1 filter
GPU Acceleration of SQL Analytics on Compressed Data
Zezhou Huang, Krystian Sakowski, Hans Lehnert +5
GPUs are uniquely suited to accelerate (SQL) analytics workloads thanks to their massive compute parallelism and High Bandwidth Memory (HBM) -- when datasets fit in the GPU HBM, pe…
Terabyte-Scale Analytics in the Blink of an Eye
Bowen Wu, Wei Cui, Carlo Curino +2
For the past two decades, the DB community has devoted substantial research to take advantage of cheap clusters of machines for distributed data analytics -- we believe that we are…
XTable in Action: Seamless Interoperability in Data Lakes
Ashvin Agrawal, Tim Brown, Anoop Johnson +4
Contemporary approaches to data management are increasingly relying on unified analytics and AI platforms to foster collaboration, interoperability, seamless access to reliable dat…
LST-Bench: Benchmarking Log-Structured Tables in the Cloud
Jesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid +6
Data processing engines increasingly leverage distributed file systems for scalable, cost-effective storage. While the Apache Parquet columnar format has become a popular choice fo…
The Tensor Data Platform: Towards an AI-centric Database System
Apurva Gandhi, Yuki Asada, Victor Fu +6
Database engines have historically absorbed many of the innovations in data processing, adding features to process graph data, XML, object oriented, and text among many others. In…
KEA: Tuning an Exabyte-Scale Data Infrastructure
Yiwen Zhu, Subru Krishnan, Konstantinos Karanasos +12
Microsoft's internal big-data infrastructure is one of the largest in the world -- with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infra…