2 papers
cs.CL2025
Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
Stefan Krsteski, Giuseppe Russo, Serina Chang +2
Surveys provide valuable insights into public opinion and behavior, but their execution is costly and slow. Large language models (LLMs) have been proposed as a scalable, low-cost…
cs.CL2025
ChatBench: From Static Benchmarks to Human-AI Evaluation
Serina Chang, Ashton Anderson, Jake M. Hofman
With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure L…