1 citations · 1 across the 6 of their papers we have counts for
3 papers · 1 filter
The Capability Frontier: Benchmarks Miss 82% of Model Performance
Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8
Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
Alexandra Yost, Shreyans Jain, Shivam Raval +6
Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We prop…
Sycophancy as compositions of Atomic Psychometric Traits
Shreyans Jain, Alexandra Yost, Amirali Abdullah
Sycophancy is a key behavioral risk in LLMs, yet is often treated as an isolated failure mode that occurs via a single causal mechanism. We instead propose modeling it as geometric…