2 papers
cs.AI2026
Counsel: A Meta-Evaluation Dataset for Agentic Tasks
Sashank Pisupati, Henry Broomfield, Eujeong Choi +5
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agen…
cs.CL2025
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…