SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
arXiv:2304.04565 · doi:10.1109/cvprw59228.2023.00536
Abstract
Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, with so many games happening simultaneously, fans cannot watch them all live. Notifications for main actions can help, but lack the engagement of live commentary, leaving fans feeling disconnected. To fulfill this need, we propose in this paper a novel task of dense video captioning focusing on the generation of textual commentaries anchored with single timestamps. To support this task, we additionally present a challenging dataset consisting of almost 37k timestamped commentaries across 715.9 hours of soccer broadcast videos. Additionally, we propose a first benchmark and baseline for this task, highlighting the difficulty of temporally anchoring commentaries yet showing the capacity to generate meaningful commentaries. By providing broadcasters with a tool to summarize the content of their video with the same level of engagement as a live game, our method could help satisfy the needs of the numerous fans who follow their team but cannot necessarily watch the live game. We believe our method has the potential to enhance the accessibility and understanding of soccer content for a wider audience, bringing the excitement of the game to more people.
References in corpus (14)
- Flamingo: a Visual Language Model for Few-Shot Learning
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Body Part-Based Representation Learning for Occluded Person Re-Identification
- SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
- DeepSportradar-v1: Computer Vision Dataset for Sports Understanding with High Quality Annotations
- Semi-Supervised Training to Improve Player and Ball Detection in Soccer
- SoccerNet 2022 Challenges Results
- Ball 3D Localization From A Single Calibrated Image
- Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- A Graph-Based Method for Soccer Action Spotting Using Unsupervised Player Classification
- Deep soccer captioning with transformer: dataset, semantics-related losses, and multi-level evaluation
- Action Spotting using Dense Detection Anchors Revisited: Submission to the SoccerNet Challenge 2022
- Going for GOAL: A Resource for Grounded Football Commentaries
Cited by in corpus (7)
- SoccerNet 2023 Challenges Results
- X-VARS: Introducing Explainability in Football Refereeing with Multi-Modal Large Language Model
- Multi-task Learning for Joint Re-identification, Team Affiliation, and Role Classification for Sports Visual Tracking
- SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries
- PLayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
- SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
- Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding