7 citations · 12 across the 2 of their papers we have counts for
4 papers
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts +15
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render ben…
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Arka Pal, Deep Karkhanis, Samuel Dooley +3
Direct Preference Optimisation (DPO) is effective at significantly improving the performance of large language models (LLMs) on downstream tasks such as reasoning, summarisation, a…
ForecastPFN: Synthetically-Trained Zero-Shot Forecasting
Samuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra +2
The vast majority of time-series forecasting approaches require a substantial training dataset. However, many real-life forecasting applications have very little initial observatio…
Giraffe: Adventures in Expanding Context Lengths in LLMs
Arka Pal, Deep Karkhanis, Manley Roberts +3
Modern large language models (LLMs) that rely on attention mechanisms are typically trained with fixed context lengths which enforce upper limits on the length of input sequences t…