7 papers
Language-Specific Gaps in AI Safety Training Datasets
Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking user…
Oasis: Hiding the Cost of Querying Parquet Files in the Datapath
Jonas Dann, Luca Tagliavini, Gustavo Alonso
Cloud-native database systems disaggregate compute and storage resources to improve cost efficiency over traditional monolithic architectures through elasticity and resource poolin…
Eiger: An Efficient Library for GPU-based Data Analytics
Bowen Wu, Marko KabiÄ, Sven Hepkema +3
GPUs have become an increasingly attractive platform for accelerating analytical workloads due to their massive parallelism and high memory bandwidth. Recent studies show that in s…
To GPU or Not to GPU: Vector Search in Relational Engines
Vasilis Mageirakos, Joel André, Marko KabiÄ +3
Vector search (VS) is now available in most database engines. However, while vector search is a common feature in AI/ML/LLMs where the dominant computing platforms are GPUs, existi…
Should I Hide My Duck in the Lake?
Jonas Dann, Gustavo Alonso
Data lakes spend a significant fraction of query execution time on scanning data from remote, disaggregated storage. Decoding alone accounts for 46% of runtime when running TPC-H d…
Cracking Vector Search Indexes
Vasilis Mageirakos, Bowen Wu, Gustavo Alonso
Retrieval Augmented Generation (RAG) uses vector databases to expand the expertise of an LLM model without having to retrain it. The idea can be applied over data lakes, leading to…