3 papers
cs.CL2026
RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
Damian Hodel, Jevin West, Aylin Caliskan
Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generalizable evaluation methods of th…
cs.AI2026
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Harshita Chopra, Kshitish Ghate, Aylin Caliskan +3
Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are unclear, impatient, or reluctant…
cs.CL2026
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
Dean Light, Michael Theologitis, Kshitish Ghate +7
Humans intuitively solve complex problems by flexibly shifting among reasoning modes: they plan, execute, revise intermediate goals, resolve ambiguity through associative judgment,…