11 citations · 18 across the 4 of their papers we have counts for
4 papers
Frontier Language Models are not Robust to Adversarial Arithmetic, or "What do I need to say so you agree 2+2=5?
C. Daniel Freeman, Laura Culp, Aaron Parisi +27
We introduce and study the problem of adversarial arithmetic, which provides a simple yet challenging testbed for language model alignment. This problem is comprised of arithmetic…
Improving Large Language Model Fine-tuning for Solving Math Problems
Yixin Liu, Avi Singh, C. Daniel Freeman +2
Despite their success in many natural language tasks, solving math problems remains a significant challenge for large language models (LLMs). A large gap exists between LLMs' pass-…
Small-scale proxies for large-scale Transformer training instabilities
Mitchell Wortsman, Peter J. Liu, Lechao Xiao +13
Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smalle…
SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu +3
Learning from human feedback has been shown to be effective at aligning language models with human preferences. Past work has often relied on Reinforcement Learning from Human Feed…