Publications (22)
K*-Means: A Parameter-free Clustering Algorithm
Louis Mahon, Mirella Lapata
Clustering is a widely used and powerful machine learning technique, but its effectiveness is often limited by the need to specify the number of clusters, k, or by relying on thres…
Robust detection of overlapping bioacoustic sound events
Louis Mahon, Benjamin Hoffman, Logan James +7
We propose a method for accurately detecting bioacoustic sound events that is robust to overlapping events, a common issue in domains such as ethology, ecology and conservation. Wh…
The Proof is in the Pudding: Using Automated Theorem Proving to Generate Cooking Recipes
Louis Mahon, Carl Vogel
This paper presents FASTFOOD, a rule-based Natural Language Generation Program for cooking recipes. Recipes are generated by using an Automated Theorem Proving procedure to select…
Minimum Description Length Clustering to Measure Meaningful Image Complexity
Louis Mahon, Thomas Lukasiewicz
Existing image complexity metrics cannot distinguish meaningful content from noise. This means that white noise images, which contain no meaningful information, are judged as highl…
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
Louis Mahon, Elliot Ford, Callum Hackett
How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs. Explainability is crucial for AI adop…
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
Dongqi Liu, Chenxi Whitehouse, Xi Yu +6
Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed…
Towards a Universal Method for Meaningful Signal Detection
Louis Mahon
It is known that human speech and certain animal vocalizations can convey meaningful content because we can decipher the content that a given utterance does convey. This paper expl…
-TCVAE: On the relationship between Disentanglement and Diversity
Cristian Meo, Louis Mahon, Anirudh Goyal +1
While disentangled representations have shown promise in generative modeling and representation learning, their downstream usefulness remains debated. Recent studies re-defined dis…
Detection-Fusion for Knowledge Graph Extraction from Videos
Taniya Das, Louis Mahon, Thomas Lukasiewicz
One of the challenging tasks in the field of video understanding is extracting semantic content from video inputs. Most existing systems use language models to describe videos in n…
Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech
Ida Szubert, Omri Abend, Nathan Schneider +4
This paper proposes a methodology for constructing such corpora of child directed speech (CDS) paired with sentential logical forms, and uses this method to create two such corpora…
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for textual descriptions or summaries that allow users to recall key plot points or get an overview without watching.…
Local Compositional Complexity: How to Detect a Human-readable Messsage
Louis Mahon
Data complexity is an important concept in the natural sciences and related areas, but lacks a rigorous and computable definition. In this paper, we focus on a particular sense of…
A Language-agnostic Model of Child Language Acquisition
Louis Mahon, Omri Abend, Uri Berger +3
This work reimplements a recent semantic bootstrapping child-language acquisition model, which was originally designed for English, and trains it to learn a new language: Hebrew. T…
Knowledge Graph Extraction from Videos
Louis Mahon, Eleonora Giunchiglia, Bowen Li +1
Nearly all existing techniques for automated video annotation (or captioning) describe videos using natural language sentences. However, this has several shortcomings: (i) it is ve…
Modelling Child Learning and Parsing of Long-range Syntactic Dependencies
Louis Mahon, Mark Johnson, Mark Steedman
This work develops a probabilistic child language acquisition model to learn a range of linguistic phenonmena, most notably long-range syntactic dependencies of the sort found in o…
A Modular Approach for Multimodal Summarization of TV Shows
Louis Mahon, Mirella Lapata
In this paper we address the task of summarizing television shows, which touches key areas in AI research: complex reasoning, multiple modalities, and long narratives. We present a…
Efficient Deep Clustering of Human Activities and How to Improve Evaluation
Louis Mahon, Thomas Lukasiewicz
There has been much recent research on human activity re\-cog\-ni\-tion (HAR), due to the proliferation of wearable sensors in watches and phones, and the advances of deep learning…
On the Existence and Behavior of Secondary Attention Sinks
Jeffrey T. H. Wong, Cheng Zhang, Louis Mahon +3
Attention sinks are tokens, often the beginning-of-sequence (BOS) token, that receive disproportionately high attention despite limited semantic relevance. In this work, we identif…
Parameter-free Video Segmentation for Vision and Language Understanding
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for adapting language models to handle video input and enable multimodal understanding. However, end-to-end models str…
Correcting Flaws in Common Disentanglement Metrics
Louis Mahon, Lei Shah, Thomas Lukasiewicz
Recent years have seen growing interest in learning disentangled representations, in which distinct features, such as size or shape, are represented by distinct neurons. Quantifyin…
Hard Regularization to Prevent Deep Online Clustering Collapse without Data Augmentation
Louis Mahon, Thomas Lukasiewicz
Online deep clustering refers to the joint use of a feature extraction network and a clustering model to assign cluster labels to each new data point or batch as it is processed. W…
Selective Pseudo-label Clustering
Louis Mahon, Thomas Lukasiewicz
Deep neural networks (DNNs) offer a means of addressing the challenging task of clustering high-dimensional data. DNNs can extract useful features, and so produce a lower dimension…