collaborators

10 papers

stat.ME2026

Tuning-Free Efficient Estimation for Multi-Source Data via Covariance-Aware Shrinkage

Wenbo Jing, Xi Chen, Yaqi Duan +2

Modern statistical learning problems often involve multiple related data sets, where learning efficiency on a target set can be improved by utilizing related source sets, while het…

cs.LG2026

LLM-Powered Virtual Population for Demand Simulation and Pricing

Chengpiao Huang, Kaizheng Wang

We develop an LLM-powered virtual population model that simulates demand for pricing decisions, in settings where products are described by rich unstructured information, such as t…

stat.ME2026

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective

Chengpiao Huang, Yuhang Wu, Kaizheng Wang

Large language models (LLMs) are increasingly used to simulate survey responses, but synthetic data can be misaligned with the human population, leading to unreliable inference. We…

stat.ML2026

Adaptive Querying with AI Persona Priors

Kaizheng Wang, Yuhang Wu, Assaf Zeevi

We study adaptive querying for learning user-dependent quantities of interest, such as responses to held-out items and psychometric indicators, within tight query budgets. Classica…

stat.ME2026

Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression

Seok-Jin Kim, Kaizheng Wang

We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A central challenge in this set…

cs.LG2026

SYN-DIGITS: A Synthetic Control Framework for Calibrated Digital Twin Simulation

Grace Jiarui Fan, Chengpiao Huang, Tianyi Peng +2

AI-based persona simulation -- often referred to as digital twin simulation -- is increasingly used for market research, recommender systems, and social sciences. Despite their fle…