6 papers
Rigorous Interpretation Is a Form of Evaluation
Isabelle Lee, Emmy Liu, Cathy Jiao +4
Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, howe…
Efficient Dataset Selection for Continual Adaptation of Generative Recommenders
Cathy Jiao, Juan Elenter, Praveen Ravichandran +7
Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impr…
An Economic Framework for Generative Engines: Advertising or Subscription?
Luyang Zhang, Cathy Jiao, Beibei Li +1
Generative Engines (GEs) such as ChatGPT and Google's AI Overviews are rapidly reshaping search economics by delivering synthesized responses that allow users to bypass third-party…
Fairshare Data Pricing via Data Valuation for Large Language Models
Luyang Zhang, Cathy Jiao, Beibei Li +1
Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with litt…
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
Cathy Jiao, Yijun Pan, Emily Xiao +6
Data attribution methods quantify the influence of training data on model outputs and are becoming increasingly relevant for a wide range of LLM research and applications, includin…
On the Feasibility of In-Context Probing for Data Attribution
Cathy Jiao, Gary Gao, Aditi Raghunathan +1
Data attribution methods are used to measure the contribution of training data towards model outputs, and have several important applications in areas such as dataset curation and…