activity
20162025
most citedERA: A Framework for Economic Resource Allocation for the Cloud

40 citations · 123 across the 10 of their papers we have counts for

collaborators
Showing cs.DBShow all

9 papers · 1 filter

cs.DB20251 cited

GPU Acceleration of SQL Analytics on Compressed Data

Zezhou Huang, Krystian Sakowski, Hans Lehnert +5

GPUs are uniquely suited to accelerate (SQL) analytics workloads thanks to their massive compute parallelism and High Bandwidth Memory (HBM) -- when datasets fit in the GPU HBM, pe…

cs.DB2025

Terabyte-Scale Analytics in the Blink of an Eye

Bowen Wu, Wei Cui, Carlo Curino +2

For the past two decades, the DB community has devoted substantial research to take advantage of cheap clusters of machines for distributed data analytics -- we believe that we are…

cs.DB20244 cited

XTable in Action: Seamless Interoperability in Data Lakes

Ashvin Agrawal, Tim Brown, Anoop Johnson +4

Contemporary approaches to data management are increasingly relying on unified analytics and AI platforms to foster collaboration, interoperability, seamless access to reliable dat…

cs.DB2023

LST-Bench: Benchmarking Log-Structured Tables in the Cloud

Jesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid +6

Data processing engines increasingly leverage distributed file systems for scalable, cost-effective storage. While the Apache Parquet columnar format has become a popular choice fo…

cs.DB20228 cited

The Tensor Data Platform: Towards an AI-centric Database System

Apurva Gandhi, Yuki Asada, Victor Fu +6

Database engines have historically absorbed many of the innovations in data processing, adding features to process graph data, XML, object oriented, and text among many others. In…

cs.DB202116 cited

KEA: Tuning an Exabyte-Scale Data Infrastructure

Yiwen Zhu, Subru Krishnan, Konstantinos Karanasos +12

Microsoft's internal big-data infrastructure is one of the largest in the world -- with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infra…