4 papers
Auditable Context-Aware HFMD Forecasting with Structured LLM Agents
Joongwon Chae, Runming Wang, Chen Xiong +5
Effective HFMD surveillance requires forecasts capturing both time-series patterns and contextual drivers such as school calendars, weather, and policy or surveillance reports. In…
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
Cathy Jiao, Yijun Pan, Emily Xiao +6
Data attribution methods quantify the influence of training data on model outputs and are becoming increasingly relevant for a wide range of LLM research and applications, includin…
Fairshare Data Pricing via Data Valuation for Large Language Models
Luyang Zhang, Cathy Jiao, Beibei Li +1
Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with litt…
On the Feasibility of In-Context Probing for Data Attribution
Cathy Jiao, Gary Gao, Aditi Raghunathan +1
Data attribution methods are used to measure the contribution of training data towards model outputs, and have several important applications in areas such as dataset curation and…