3 papers
cs.LG2026
Interpreting and Steering LLM Agents for Social Simulations
Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan +1
Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. Howe…
cs.CL2026
ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge
Shreyas Krishnan, Serina Chang, Abhishek Nagaraj
We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utili…
cs.CY2026
CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong +1
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives…