6 papers
CEO-Bench: Can Agents Play the Long Game?
Haozhe Chen, Karthik Narasimhan, Zhuang Liu
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a comb…
Human-like Navigation in a World Built for Humans
Bhargav Chandaka, Gloria X. Wang, Haozhe Chen +3
When navigating in a man-made environment they haven't visited before--like an office building--humans employ behaviors such as reading signs and asking others for directions. Thes…
Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
Olivier Toubia, George Z. Gui, Tianyi Peng +3
LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital ex…
Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
Thomson Yen, Andrew Wei Tung Siah, Haozhe Chen +3
Careful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, maki…
LLM Generated Persona is a Promise with a Catch
Ang Li, Haozhe Chen, Hongseok Namkoong +1
The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Pers…
QGym: Scalable Simulation and Benchmarking of Queuing Network Controllers
Haozhe Chen, Ang Li, Ethan Che +3
Queuing network control determines the allocation of scarce resources to manage congestion, a fundamental problem in manufacturing, communications, and healthcare. Compared to stan…