1 citations · 1 across the 9 of their papers we have counts for
4 papers · 1 filter
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
Ahmed Hassoon, Mark Dredze
Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to…
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang, Fan Bai, Zhongyan Luo +9
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
Fatima Jahara, Mark Dredze, Sharon Levy
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation…
How to Interpret Agent Behavior
Jie Gao, Kaiser Sun, Jen-tse Huang +8
Autonomous agents such as Claude Code and Codex now operate for hours or even days. Understanding their runtime behavior has become critical for downstream tasks such as diagnosing…