collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

Seven simple steps for log analysis in AI systems

Magda Dubois, Ekin Zorer, Maia Hamin +17

AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…

cs.AI2026

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

Sunishchal Dev, Andrew Sloan, Joshua Kavner +2

We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed…

cs.AI2026

ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs

Rohan Subramanian Thomas, Shikhar Shiromani, Abdullah Chaudhry +4

Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and model…

cs.AI2025

Emergent Persuasion: Will LLMs Persuade Without Being Prompted?

Vincent Chang, Thee Ho, Sunishchal Dev +4

With the wide-scale adoption of conversational AI systems, AI are now able to exert unprecedented influence on human opinion and beliefs. Recent work has shown that many Large Lang…

cs.AI2025

AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI

Manik Rana, Calissa Man, Anotida Expected Msiiwa +5

Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce Agen…

cs.AI2025

Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games

Chris Su, Harrison Li, Matheus Marques +3

Recent work reports that Large Reasoning Models (LRMs) undergo a collapse in performance on solving puzzles beyond certain perplexity thresholds. In subsequent discourse, questions…