Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LLM Powered Social Digital Twins: A Framework for Simulating Population Behavioral Response to Policy Interventions
Fatima Koaik, Aayush Gupta, Farahan Raza Sheikh
Predicting how populations respond to policy interventions is a fundamental challenge in computational social science and public policy. Traditional approaches rely on aggregate st…
cs.AI2026
ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
Aayush Gupta
Existing benchmarks for tool-using LLM agents primarily report single-run success rates and miss reliability properties required in production. We introduce \textbf{ReliabilityBenc…