2 papers
cs.AI2026
Counsel: A Meta-Evaluation Dataset for Agentic Tasks
Sashank Pisupati, Henry Broomfield, Eujeong Choi +5
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agen…
cs.CL2025
Atla Selene Mini: A General Purpose Evaluation Model
Andrei Alexandru, Antonia Calvi, Henry Broomfield +9
We introduce Atla Selene Mini, a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini is a general-purpose evaluator that outperforms the best SLMJs and GPT-4o-mini…