collaborators

6 papers

cs.CY2026

Rigorous Interpretation Is a Form of Evaluation

Isabelle Lee, Emmy Liu, Cathy Jiao +4

Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, howe…

cs.IR2026

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders

Cathy Jiao, Juan Elenter, Praveen Ravichandran +7

Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impr…

cs.GT2026

An Economic Framework for Generative Engines: Advertising or Subscription?

Luyang Zhang, Cathy Jiao, Beibei Li +1

Generative Engines (GEs) such as ChatGPT and Google's AI Overviews are rapidly reshaping search economics by delivering synthesized responses that allow users to bypass third-party…

cs.GT2025

Fairshare Data Pricing via Data Valuation for Large Language Models

Luyang Zhang, Cathy Jiao, Beibei Li +1

Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with litt…

cs.CL2025

DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models

Cathy Jiao, Yijun Pan, Emily Xiao +6

Data attribution methods quantify the influence of training data on model outputs and are becoming increasingly relevant for a wide range of LLM research and applications, includin…

cs.CL2025

On the Feasibility of In-Context Probing for Data Attribution

Cathy Jiao, Gary Gao, Aditi Raghunathan +1

Data attribution methods are used to measure the contribution of training data towards model outputs, and have several important applications in areas such as dataset curation and…