papers

Publications (37)

cs.CL2026

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

Łukasz Borchmann, Jordy Van Landeghem, Michał Turski +12

Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoni…

math.CA2014

Structural identifiability of viscoelastic mechanical systems

Adam Mahdi, Nicolette Meshkat, Seth Sullivant

We solve the local and global structural identifiability problems for viscoelastic mechanical models represented by networks of springs and dashpots. We propose a very simple chara…

cs.CL2024

Do Large Language Models have Shared Weaknesses in Medical Question Answering?

Andrew M. Bean, Karolina Korgul, Felix Krones +2

Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the u…

math.DS2015

A hybrid symbolic-numerical approach to the center-focus problem

Adam Mahdi, Claudio Pessoa, Jonathan D. Hauenstein

We propose a new hybrid symbolic-numerical approach to the center-focus problem. The method allowed us to obtain center conditions for a three-dimensional system of differential eq…

physics.med-ph2018

Bayesian approach to uncertainty quantification for cerebral autoregulation index

Kevin P. O'Keeffe, Adam Mahdi

Cerebral autoregulation refers to the brain's ability to maintain cerebral blood flow at an approximately constant level, despite changes in arterial blood pressure. The performanc…

cs.CL2024

Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering

Yushi Yang, Andrew M. Bean, Robert McCraith +1

Fine-tuning Large Language Models (LLMs) incurs considerable training costs, driving the need for data-efficient training with optimised data ordering. Human-inspired strategies of…

cs.LG2025

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

Yushi Yang, Filip Sondej, Harry Mayne +2

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of alg…

cs.LG2024

Multimodal deep learning approach to predicting neurological recovery from coma after cardiac arrest

Felix H. Krones, Ben Walker, Guy Parsons +2

This work showcases our team's (The BEEGees) contributions to the 2023 George B. Moody PhysioNet Challenge. The aim was to predict neurological recovery from coma following cardiac…

math.OC2016

Bayesian inference in non-Markovian state-space models with applications to fractional order systems

Pierre E. Jacob, S. M. Mahdi Alavi, Adam Mahdi +2

Battery impedance spectroscopy models are given by fractional order (FO) differential equations. In the discrete-time domain, they give rise to state-space models where the latent…

cs.LG2026

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…

math.OC2015

Structural Identifiability Analysis of Fractional Order Models with Applications in Battery Systems

S. M. Mahdi Alavi, Adam Mahdi, Pierre E. Jacob +2

This paper presents a method for structural identifiability analysis of fractional order systems by using the coefficient mapping concept to determine whether the model parameters…

cs.CL2026

RepSelect: Robust LLM Unlearning via Representation Selectivity

Filip Sondej, Yushi Yang, Adam Mahdi

Making large language models (LLMs) deeply forget specific knowledge and values without sacrificing general capabilities remains a central challenge in unlearning. Current methods…

cs.LG2024

Combining Hough Transform and Deep Learning Approaches to Reconstruct ECG Signals From Printouts

Felix Krones, Ben Walker, Terry Lyons +1

This work presents our team's (SignalSavants) winning contribution to the 2024 George B. Moody PhysioNet Challenge. The Challenge had two goals: reconstruct ECG signals from printo…

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.LG2025

LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

Harry Mayne, Ryan Othniel Kearns, Yushi Yang +4

To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated co…

cs.LG2024

Can sparse autoencoders be used to decompose and interpret steering vectors?

Harry Mayne, Yushi Yang, Adam Mahdi

Steering vectors are a promising approach to control the behaviour of large language models. However, their underlying mechanisms remain poorly understood. While sparse autoencoder…

cs.CL2026

Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning

Yushi Yang, Shreyansh Padarha, Sarah Ball +2

Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…

cs.AI2026

A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

Harry Mayne, Justin Singh Kang, Dewi Gould +3

LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithful…

cs.CY2024

Large language models can help boost food production, but be mindful of their risks

Djavan De Clercq, Elias Nehring, Harry Mayne +1

Coverage of ChatGPT-style large language models (LLMs) in the media has focused on their eye-catching achievements, including solving advanced mathematical problems and reaching ex…

q-bio.QM2017

Increased blood pressure variability upon standing up improves reproducibility of cerebral autoregulation indices

Adam Mahdi, Dragana Nikolic, Anthony A. Birch +4

Dynamic cerebral autoregulation, that is the transient response of cerebral blood flow to changes in arterial blood pressure, is currently assessed using a variety of different tim…

q-bio.QM2020

Sensitivity analysis methods in the biomedical sciences

George Qian, Adam Mahdi

Sensitivity analysis is an important part of a mathematical modeller's toolbox for model analysis. In this review paper, we describe the most frequently used sensitivity techniques…

cs.CL2026

LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

Jude Khouja, Lingyi Yang, Karolina Korgul +6

Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…

cs.IR2026

Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews

Shreyansh Padarha, Ryan Othniel Kearns, Tristan Naidoo +13

Systematic literature reviews (SLRs) are a demanding and high-stakes form of scientific knowledge synthesis that remains underspecified as an evaluation setting for large language…

q-bio.TO2015

Mathematical model of the interaction between baroreflex and cerebral autoregulation

Adam Mahdi, Mette S. Olufsen, Stephen J. Payne

Baroreflex (BR) and cerebral autoregulation (CA) are two important mechanisms regulating blood pressure and flow. However, the functional relationship between BR and CA in humans i…

cs.LG2023

Dual Bayesian ResNet: A Deep Learning Approach to Heart Murmur Detection

Benjamin Walker, Felix Krones, Ivan Kiskin +3

This study presents our team PathToMyHeart's contribution to the George B. Moody PhysioNet Challenge 2022. Two models are implemented. The first model is a Dual Bayesian ResNet (DB…

cs.HC2025

Clinical knowledge in LLMs does not translate to human interactions

Andrew M. Bean, Rebecca Payne, Guy Parsons +8

Global healthcare providers are exploring use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing e…

cs.LG2024

Review of multimodal machine learning approaches in healthcare

Felix Krones, Umar Marikkar, Guy Parsons +2

Machine learning methods in healthcare have traditionally focused on using data from a single modality, limiting their ability to effectively replicate the clinical practice of int…

math-ph2014

Integrability of the Hide--Skeldon--Acheson dynamo

Adam Mahdi, Claudia Valls

In this work we consider the Hide-Skeldon-Acheson dynamo model \[ \dot x=x(y-1)-βz, \quad \dot y =α(1-x^2)-κy, \quad \dot z =x-λz, \] where and are parameters.…

cs.AI2024

Evaluating the role of `Constitutions' for learning from AI feedback

Saskia Redgate, Andrew M. Bean, Adam Mahdi

The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on…

cs.CL2024

Improving In-Context Learning with Small Language Model Ensembles

M. Mehdi Mojarradi, Lingyi Yang, Robert McCraith +1

Large language models (LLMs) have shown impressive capabilities across various tasks, but their performance on domain-specific tasks remains limited. While methods like retrieval a…

cs.LG2024

Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?

Harry Mayne, Guy Parsons, Adam Mahdi

The use of unsupervised learning to identify patient subgroups has emerged as a potentially promising direction to improve the efficiency of Intensive Care Units (ICUs). By identif…

cs.CV2023

LT-ViT: A Vision Transformer for multi-label Chest X-ray classification

Umar Marikkar, Sara Atito, Muhammad Awais +1

Vision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). Howev…

stat.ME2022

On automatic calibration of the SIRD epidemiological model for COVID-19 data in Poland

Piotr Błaszczyk, Konrad Klimczak, Adam Mahdi +4

We propose a novel methodology for estimating the epidemiological parameters of a modified SIRD model (acronym of Susceptible, Infected, Recovered and Deceased individuals) and per…

cs.HC2026

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7

Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, howeve…

cs.LG2024

Feasibility of machine learning-based rice yield prediction in India at the district level using climate reanalysis data

Djavan De Clercq, Adam Mahdi

Yield forecasting, the science of predicting agricultural productivity before the crop harvest occurs, helps a wide range of stakeholders make better decisions around agricultural…

cs.CL2025

Framing Migration: A Computational Analysis of UK Parliamentary Discourse

Vahid Ghafouri, Robert McNeil, Teodor Yankov +4

We present a large-scale computational analysis of migration-related discourse in UK parliamentary debates spanning over 75 years and compare it with US congressional discourse. Us…

q-bio.TO2016

Effects of non-physiological blood pressure artefacts on measures of cerebral autoregulation

Adam Mahdi, Erica Rutter, Stephen J. Payne

Cerebral autoregulation refers to regulation mechanisms that aim to maintain cerebral blood flow approximately constant. It is often assessed by autoregulation index (ARI), which u…