4 citations · 4 across the 5 of their papers we have counts for
4 papers · 1 filter
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…
Nevermind: Instruction Override and Moderation in Large Language Models
Edward Kim
Given the impressive capabilities of recent Large Language Models (LLMs), we investigate and benchmark the most popular proprietary and different sized open source models on the ta…
Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis +2
In Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large La…
Which Prompts Make The Difference? Data Prioritization For Efficient Human LLM Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis +2
Human evaluation is increasingly critical for assessing large language models, capturing linguistic nuances, and reflecting user preferences more accurately than traditional automa…