16 citations · 43 across the 3 of their papers we have counts for
6 papers
Optimizing open-domain question answering with graph-based retrieval augmented generation
Joyce Cahoon, Prerna Singh, Nick Litombe +6
In this work, we benchmark various graph-based retrieval-augmented generation (RAG) systems across a broad spectrum of query types, including OLTP-style (fact-based) and OLAP-style…
KEA: Tuning an Exabyte-Scale Data Infrastructure
Yiwen Zhu, Subru Krishnan, Konstantinos Karanasos +12
Microsoft's internal big-data infrastructure is one of the largest in the world -- with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infra…
MLOS: An Infrastructure for Automated Software Performance Engineering
Carlo Curino, Neha Godwal, Brian Kroth +9
Developing modern systems software is a complex task that combines business logic programming and Software Performance Engineering (SPE). The later is an experimental and labor-int…
Data Science through the looking glass and what we found there
Fotis Psallidas, Yiwen Zhu, Bojan Karlas +8
The recent success of machine learning (ML) has led to an explosive growth both in terms of new systems and algorithms built in industry and academia, and new applications built by…
Cloudy with high chance of DBMS: A 10-year prediction for Enterprise-Grade ML
Ashvin Agrawal, Rony Chatterjee, Carlo Curino +19
Machine learning (ML) has proven itself in high-value web applications such as search ranking and is emerging as a powerful tool in a much broader range of enterprise scenarios inc…
Griffon: Reasoning about Job Anomalies with Unlabeled Data in Cloud-based Platforms
Liqun Shao, Yiwen Zhu, Abhiram Eswaran +9
Microsoft's internal big data analytics platform is comprised of hundreds of thousands of machines, serving over half a million jobs daily, from thousands of users. The majority of…