activity
20222026
most citedComputing Rule-Based Explanations of Machine Learning Classifiers using Knowledge Graphs

3 citations · 3 across the 17 of their papers we have counts for

collaborators

18 papers

cs.AI2026

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Lin Shi, Haowei Lin, Zixuan Zhu +123

Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a…

cs.CR2026

AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis +6

LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pr…

cs.CL2026

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…

cs.CL2026

The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans

Odysseas S. Chlapanis, Orfeas Menis Mastromichalakis, Christos H. Papadimitriou

Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social con…

cs.CL2026

Explain the Flag: Contextualizing Hate Speech Beyond Censorship

Jason Liartis, Eirini Kaldeli, Lambrini Gyftokosta +2

Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are widely used, most focus on cen…

cs.SE2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…