4 papers
Which Pieces Does Unigram Tokenization Really Need?
Sander Land, Yuval Pinter
The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegance, its implementation in practi…
RewardBench 2: Advancing Reward Model Evaluation
Saumya Malik, Valentina Pyatkin, Sander Land +4
Reward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target for optimization across instruc…
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
Zhengyan Shi, Sander Land, Acyr Locatelli +2
Direct Alignment Algorithms (DAAs), such as Direct Preference Optimisation (DPO) and Identity Preference Optimisation (IPO), have emerged as alternatives to online Reinforcement Le…