10 papers
Tuning-Free Efficient Estimation for Multi-Source Data via Covariance-Aware Shrinkage
Wenbo Jing, Xi Chen, Yaqi Duan +2
Modern statistical learning problems often involve multiple related data sets, where learning efficiency on a target set can be improved by utilizing related source sets, while het…
LLM-Powered Virtual Population for Demand Simulation and Pricing
Chengpiao Huang, Kaizheng Wang
We develop an LLM-powered virtual population model that simulates demand for pricing decisions, in settings where products are described by rich unstructured information, such as t…
How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective
Chengpiao Huang, Yuhang Wu, Kaizheng Wang
Large language models (LLMs) are increasingly used to simulate survey responses, but synthetic data can be misaligned with the human population, leading to unreliable inference. We…
Adaptive Querying with AI Persona Priors
Kaizheng Wang, Yuhang Wu, Assaf Zeevi
We study adaptive querying for learning user-dependent quantities of interest, such as responses to held-out items and psychometric indicators, within tight query budgets. Classica…
Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression
Seok-Jin Kim, Kaizheng Wang
We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A central challenge in this set…
SYN-DIGITS: A Synthetic Control Framework for Calibrated Digital Twin Simulation
Grace Jiarui Fan, Chengpiao Huang, Tianyi Peng +2
AI-based persona simulation -- often referred to as digital twin simulation -- is increasingly used for market research, recommender systems, and social sciences. Despite their fle…