3 citations · 3 across the 17 of their papers we have counts for
18 papers
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Lin Shi, Haowei Lin, Zixuan Zhu +123
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a…
AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories
Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis +6
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pr…
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
Odysseas S. Chlapanis, Orfeas Menis Mastromichalakis, Christos H. Papadimitriou
Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social con…
Explain the Flag: Contextualizing Hate Speech Beyond Censorship
Jason Liartis, Eirini Kaldeli, Lambrini Gyftokosta +2
Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are widely used, most focus on cen…
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…