4 papers
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
Dongqi Liu, Chenxi Whitehouse, Xi Yu +6
Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed…
K*-Means: A Parameter-free Clustering Algorithm
Louis Mahon, Mirella Lapata
Clustering is a widely used and powerful machine learning technique, but its effectiveness is often limited by the need to specify the number of clusters, k, or by relying on thres…
Parameter-free Video Segmentation for Vision and Language Understanding
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for adapting language models to handle video input and enable multimodal understanding. However, end-to-end models str…
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
Louis Mahon, Mirella Lapata
The proliferation of creative video content has driven demand for textual descriptions or summaries that allow users to recall key plot points or get an overview without watching.…