Publications (9)
TempTest: Local Normalization Distortion and the Detection of Machine-generated Text
Tom Kempton, Stuart Burrell, Connor Cheverall
Existing methods for the zero-shot detection of machine-generated text are dominated by three statistical quantities: log-likelihood, log-rank, and entropy. As language models mimi…
DMAP: A Distribution Map for Text
Tom Kempton, Julia Rozanova, Parameswaran Kamalaruban +5
Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Ext…
Locally Differentially Private Embedding Models in Distributed Fraud Prevention Systems
Iker Perez, Jason Wong, Piotr Skalski +4
Global financial crime activity is driving demand for machine learning solutions in fraud prevention. However, prevention systems are commonly serviced to financial institutions in…
Towards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences
Piotr Skalski, David Sutton, Stuart Burrell +2
Machine learning models underpin many modern financial systems for use cases such as fraud detection and churn prediction. Most are based on supervised learning with hand-engineere…
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
Tom Kempton, Stuart Burrell
Advances in hardware and language model architecture have spurred a revolution in natural language generation. However, autoregressive models compute probability distributions over…
Emergent Bias and Fairness in Multi-Agent Decision Systems
Maeve Madigan, Parameswaran Kamalaruban, Glenn Moynihan +3
Multi-agent systems have demonstrated the ability to improve performance on a variety of predictive tasks by leveraging collaborative decision making. However, the lack of effectiv…
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
Parameswaran Kamalaruban, Yulu Pi, Stuart Burrell +4
Ensuring fairness in transaction fraud detection models is vital due to the potential harms and legal implications of biased decision-making. Despite extensive research on algorith…
Fairness-Aware Low-Rank Adaptation Under Demographic Privacy Constraints
Parameswaran Kamalaruban, Mark Anderson, Stuart Burrell +3
Pre-trained foundation models can be adapted for specific tasks using Low-Rank Adaptation (LoRA). However, the fairness properties of these adapted classifiers remain underexplored…
Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
Tom Kempton, Viktor Drobnyi, Maeve Madigan +1
The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploi…