23 citations · 64 across the 19 of their papers we have counts for
5 papers · 1 filter
Rethinking Dataset Discovery with DataScout
Rachel Lin, Bhavya Chopra, Wenjing Lin +3
Dataset Search -- the process of finding appropriate datasets for a given task -- remains a critical yet under-explored challenge in data science workflows. Assessing dataset suita…
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann +2
Due to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs) are increasingly being used to assist humans in evaluating L…
"We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning
Shreya Shankar, Rolando Garcia, Joseph M Hellerstein +1
Organizations rely on machine learning engineers (MLEs) to deploy models and maintain ML pipelines in production. Due to models' extensive reliance on fresh data, the operationaliz…
Understanding Workers, Developing Effective Tasks, and Enhancing Marketplace Dynamics: A Study of a Large Crowdsourcing Marketplace
Ayush Jain, Akash Das Sarma, Aditya Parameswaran +1
We conduct an experimental analysis of a dataset comprising over 27 million microtasks performed by over 70,000 workers issued to a large crowdsourcing marketplace between 2012-201…
Optimizing Open-Ended Crowdsourcing: The Next Frontier in Crowdsourced Data Management
Aditya Parameswaran, Akash Das Sarma, Vipul Venkataraman
Crowdsourcing is the primary means to generate training data at scale, and when combined with sophisticated machine learning algorithms, crowdsourcing is an enabler for a variety o…