2 citations · 2 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 6 cited
Understanding the Dataset Practitioners Behind Large Language Model Development
Crystal Qian, Emily Reif, Minsuk Kahng
As large language models (LLMs) become more advanced and impactful, it is increasingly important to scrutinize the data that they rely upon and produce. What is it to be a dataset…
cs.CL2024
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
Emily Reif, Crystal Qian, James Wexler +1
Making sense of unstructured text datasets is perennially difficult, yet increasingly relevant with Large Language Models. Data workers often rely on dataset summaries, especially…