2 papers
cs.AI2026
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
Jie Ruan, Inderjeet Nair, Amy Liu +3
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instr…
cs.CL2025
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
Jie Ruan, Inderjeet Nair, Shuyang Cao +14
This paper introduces ExpertLongBench, an expert-level benchmark containing 11 tasks from 9 domains that reflect realistic expert workflows and applications. Beyond question answer…