activity
20242026
collaborators

6 papers

cs.LG2026

MADE: Benchmark Environments for Closed-Loop Materials Discovery

Shreshth A Malik, Tiarnan Doherty, Panagiotis Tigas +4

Existing benchmarks for computational materials discovery primarily evaluate static predictive tasks or isolated computational sub-tasks. While valuable, these evaluations neglect…

cs.LG2025

Scaling Up Active Testing to Large Language Models

Gabrielle Berrada, Jannik Kossen, Freddie Bickford Smith +3

Active testing enables label-efficient evaluation of predictive models through careful data acquisition, but it can pose a significant computational cost. We identify cost-saving m…

cs.AI2025

Robin: A multi-agent system for automating scientific discovery

Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener +7

Scientific discovery is driven by the iterative process of background research, hypothesis generation, experimentation, and data analysis. Despite recent advancements in applying a…

cs.CL2024

Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy

Benedict Aaron Tjandra, Muhammed Razzak, Jannik Kossen +2

Large Language Models (LLMs) are known to hallucinate, whereby they generate plausible but inaccurate text. This phenomenon poses significant risks in critical applications, such a…

cs.CL2024

Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Jannik Kossen, Jiatong Han, Muhammed Razzak +3

We propose semantic entropy probes (SEPs), a cheap and reliable method for uncertainty quantification in Large Language Models (LLMs). Hallucinations, which are plausible-sounding…

cs.LG2024

The Benefits and Risks of Transductive Approaches for AI Fairness

Muhammed Razzak, Andreas Kirsch, Yarin Gal

Recently, transductive learning methods, which leverage holdout sets during training, have gained popularity for their potential to improve speed, accuracy, and fairness in machine…