5 papers
BRIDGE: Predicting Human Task Completion Time From Model Performance
Fengyuan Liu, Jay Gala, Nilaksh +3
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on d…
Operationalising the Superficial Alignment Hypothesis via Task Complexity
Tomás Vergara-Browne, Darshan Patil, Ivan Titov +3
The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledg…
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
Mridul Sharma, Adeetya Patel, Zaneta D' Souza +3
Despite their widespread applications, Large Language Models (LLMs) often struggle to express uncertainty, posing a challenge for reliable deployment in high stakes and safety crit…
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
Le Zhang, Bo Wang, Xipeng Qiu +2
We present REARANK, a large language model (LLM)-based listwise reasoning reranking agent. REARANK explicitly reasons before reranking, significantly improving both performance and…
How to Get Your LLM to Generate Challenging Problems for Evaluation
Arkil Patel, Siva Reddy, Dzmitry Bahdanau
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impractica…