8 citations · 8 across the 2 of their papers we have counts for
3 papers
Tab-Shapley: Identifying Top-k Tabular Data Quality Insights
Manisha Padala, Lokesh Nagalapatti, Atharv Tyagi +2
We present an unsupervised method for aggregating anomalies in tabular datasets by identifying the top-k tabular data quality insights. Each insight consists of a set of anomalous…
Towards Optimizing the Costs of LLM Usage
Shivanshu Shekhar, Tanishq Dubey, Koyel Mukherjee +3
Generative AI and LLMs in particular are heavily used nowadays for various document processing tasks such as question answering and summarization. However, different LLMs come with…
R2D2: Reducing Redundancy and Duplication in Data Lakes
Raunak Shah, Koyel Mukherjee, Atharv Tyagi +4
Enterprise data lakes often suffer from substantial amounts of duplicate and redundant data, with data volumes ranging from terabytes to petabytes. This leads to both increased sto…