1 citations · 1 across the 8 of their papers we have counts for
29 papers
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
Ahmed Hassoon, Mark Dredze
Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to…
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Ahmed Hassoon, Mark Dredze
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these…
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Zhen Wang, Fan Bai, Zhongyan Luo +9
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery r…
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
Fatima Jahara, Mark Dredze, Sharon Levy
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation…
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
Jen-tse Huang, Chang Chen, Shiyang Lai +3
Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language…
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4
Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematica…