Publications (37)
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Åukasz Borchmann, Jordy Van Landeghem, MichaÅ Turski +12
Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoni…
Structural identifiability of viscoelastic mechanical systems
Adam Mahdi, Nicolette Meshkat, Seth Sullivant
We solve the local and global structural identifiability problems for viscoelastic mechanical models represented by networks of springs and dashpots. We propose a very simple chara…
Do Large Language Models have Shared Weaknesses in Medical Question Answering?
Andrew M. Bean, Karolina Korgul, Felix Krones +2
Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the u…
A hybrid symbolic-numerical approach to the center-focus problem
Adam Mahdi, Claudio Pessoa, Jonathan D. Hauenstein
We propose a new hybrid symbolic-numerical approach to the center-focus problem. The method allowed us to obtain center conditions for a three-dimensional system of differential eq…
Bayesian approach to uncertainty quantification for cerebral autoregulation index
Kevin P. O'Keeffe, Adam Mahdi
Cerebral autoregulation refers to the brain's ability to maintain cerebral blood flow at an approximately constant level, despite changes in arterial blood pressure. The performanc…
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
Yushi Yang, Andrew M. Bean, Robert McCraith +1
Fine-tuning Large Language Models (LLMs) incurs considerable training costs, driving the need for data-efficient training with optimised data ordering. Human-inspired strategies of…
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
Yushi Yang, Filip Sondej, Harry Mayne +2
Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of alg…
Multimodal deep learning approach to predicting neurological recovery from coma after cardiac arrest
Felix H. Krones, Ben Walker, Guy Parsons +2
This work showcases our team's (The BEEGees) contributions to the 2023 George B. Moody PhysioNet Challenge. The aim was to predict neurological recovery from coma following cardiac…
Bayesian inference in non-Markovian state-space models with applications to fractional order systems
Pierre E. Jacob, S. M. Mahdi Alavi, Adam Mahdi +2
Battery impedance spectroscopy models are given by fractional order (FO) differential equations. In the discrete-time domain, they give rise to state-space models where the latent…
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…
Structural Identifiability Analysis of Fractional Order Models with Applications in Battery Systems
S. M. Mahdi Alavi, Adam Mahdi, Pierre E. Jacob +2
This paper presents a method for structural identifiability analysis of fractional order systems by using the coefficient mapping concept to determine whether the model parameters…
RepSelect: Robust LLM Unlearning via Representation Selectivity
Filip Sondej, Yushi Yang, Adam Mahdi
Making large language models (LLMs) deeply forget specific knowledge and values without sacrificing general capabilities remains a central challenge in unlearning. Current methods…
Combining Hough Transform and Deep Learning Approaches to Reconstruct ECG Signals From Printouts
Felix Krones, Ben Walker, Terry Lyons +1
This work presents our team's (SignalSavants) winning contribution to the 2024 George B. Moody PhysioNet Challenge. The Challenge had two goals: reconstruct ECG signals from printo…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
Harry Mayne, Ryan Othniel Kearns, Yushi Yang +4
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated co…
Can sparse autoencoders be used to decompose and interpret steering vectors?
Harry Mayne, Yushi Yang, Adam Mahdi
Steering vectors are a promising approach to control the behaviour of large language models. However, their underlying mechanisms remain poorly understood. While sparse autoencoder…
Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
Yushi Yang, Shreyansh Padarha, Sarah Ball +2
Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
Harry Mayne, Justin Singh Kang, Dewi Gould +3
LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithful…
Large language models can help boost food production, but be mindful of their risks
Djavan De Clercq, Elias Nehring, Harry Mayne +1
Coverage of ChatGPT-style large language models (LLMs) in the media has focused on their eye-catching achievements, including solving advanced mathematical problems and reaching ex…
Increased blood pressure variability upon standing up improves reproducibility of cerebral autoregulation indices
Adam Mahdi, Dragana Nikolic, Anthony A. Birch +4
Dynamic cerebral autoregulation, that is the transient response of cerebral blood flow to changes in arterial blood pressure, is currently assessed using a variety of different tim…
Sensitivity analysis methods in the biomedical sciences
George Qian, Adam Mahdi
Sensitivity analysis is an important part of a mathematical modeller's toolbox for model analysis. In this review paper, we describe the most frequently used sensitivity techniques…
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
Jude Khouja, Lingyi Yang, Karolina Korgul +6
Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…
Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews
Shreyansh Padarha, Ryan Othniel Kearns, Tristan Naidoo +13
Systematic literature reviews (SLRs) are a demanding and high-stakes form of scientific knowledge synthesis that remains underspecified as an evaluation setting for large language…
Mathematical model of the interaction between baroreflex and cerebral autoregulation
Adam Mahdi, Mette S. Olufsen, Stephen J. Payne
Baroreflex (BR) and cerebral autoregulation (CA) are two important mechanisms regulating blood pressure and flow. However, the functional relationship between BR and CA in humans i…
Dual Bayesian ResNet: A Deep Learning Approach to Heart Murmur Detection
Benjamin Walker, Felix Krones, Ivan Kiskin +3
This study presents our team PathToMyHeart's contribution to the George B. Moody PhysioNet Challenge 2022. Two models are implemented. The first model is a Dual Bayesian ResNet (DB…
Clinical knowledge in LLMs does not translate to human interactions
Andrew M. Bean, Rebecca Payne, Guy Parsons +8
Global healthcare providers are exploring use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing e…
Review of multimodal machine learning approaches in healthcare
Felix Krones, Umar Marikkar, Guy Parsons +2
Machine learning methods in healthcare have traditionally focused on using data from a single modality, limiting their ability to effectively replicate the clinical practice of int…
Integrability of the Hide--Skeldon--Acheson dynamo
Adam Mahdi, Claudia Valls
In this work we consider the Hide-Skeldon-Acheson dynamo model \[ \dot x=x(y-1)-βz, \quad \dot y =α(1-x^2)-κy, \quad \dot z =x-λz, \] where and are parameters.…
Evaluating the role of `Constitutions' for learning from AI feedback
Saskia Redgate, Andrew M. Bean, Adam Mahdi
The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on…
Improving In-Context Learning with Small Language Model Ensembles
M. Mehdi Mojarradi, Lingyi Yang, Robert McCraith +1
Large language models (LLMs) have shown impressive capabilities across various tasks, but their performance on domain-specific tasks remains limited. While methods like retrieval a…
Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?
Harry Mayne, Guy Parsons, Adam Mahdi
The use of unsupervised learning to identify patient subgroups has emerged as a potentially promising direction to improve the efficiency of Intensive Care Units (ICUs). By identif…
LT-ViT: A Vision Transformer for multi-label Chest X-ray classification
Umar Marikkar, Sara Atito, Muhammad Awais +1
Vision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). Howev…
On automatic calibration of the SIRD epidemiological model for COVID-19 data in Poland
Piotr BÅaszczyk, Konrad Klimczak, Adam Mahdi +4
We propose a novel methodology for estimating the epidemiological parameters of a modified SIRD model (acronym of Susceptible, Infected, Recovered and Deceased individuals) and per…
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7
Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, howeve…
Feasibility of machine learning-based rice yield prediction in India at the district level using climate reanalysis data
Djavan De Clercq, Adam Mahdi
Yield forecasting, the science of predicting agricultural productivity before the crop harvest occurs, helps a wide range of stakeholders make better decisions around agricultural…
Framing Migration: A Computational Analysis of UK Parliamentary Discourse
Vahid Ghafouri, Robert McNeil, Teodor Yankov +4
We present a large-scale computational analysis of migration-related discourse in UK parliamentary debates spanning over 75 years and compare it with US congressional discourse. Us…
Effects of non-physiological blood pressure artefacts on measures of cerebral autoregulation
Adam Mahdi, Erica Rutter, Stephen J. Payne
Cerebral autoregulation refers to regulation mechanisms that aim to maintain cerebral blood flow approximately constant. It is often assessed by autoregulation index (ARI), which u…