3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.AI2026★ 3 cited
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
Longling Geng, Andy Ouyang, Theodore Wu +10
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, aban…
cs.CV2026
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
Jing Wu, Daphne Barretto, Yiye Chen +4
Long-horizon, repetitive workflows are common in professional settings, such as processing expense reports from receipts and entering student grades from exam papers. These tasks a…