8 papers · 1 filter
Seven simple steps for log analysis in AI systems
Magda Dubois, Ekin Zorer, Maia Hamin +17
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
Sunishchal Dev, Andrew Sloan, Joshua Kavner +2
We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed…
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
Rohan Subramanian Thomas, Shikhar Shiromani, Abdullah Chaudhry +4
Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and model…
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
Vincent Chang, Thee Ho, Sunishchal Dev +4
With the wide-scale adoption of conversational AI systems, AI are now able to exert unprecedented influence on human opinion and beliefs. Recent work has shown that many Large Lang…
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
Manik Rana, Calissa Man, Anotida Expected Msiiwa +5
Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce Agen…
Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
Chris Su, Harrison Li, Matheus Marques +3
Recent work reports that Large Reasoning Models (LRMs) undergo a collapse in performance on solving puzzles beyond certain perplexity thresholds. In subsequent discourse, questions…