5 papers
DGME-T: Directional Grid Motion Encoding for Transformer-Based Historical Camera Movement Classification
Tingyu Lin, Armin Dadras, Florian Kleber +1
Camera movement classification (CMC) models trained on contemporary, high-quality footage often degrade when applied to archival film, where noise, missing frames, and low contrast…
ClapperText: A Benchmark for Text Recognition in Low-Resource Archival Documents
Tingyu Lin, Marco Peer, Florian Kleber +1
This paper presents ClapperText, a benchmark dataset for handwritten and printed text recognition in visually degraded and low-resource settings. The dataset is derived from 127 Wo…
Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models
Tingyu Lin, Armin Dadras, Florian Kleber +1
Camera movement conveys spatial and narrative information essential for understanding video content. While recent camera movement classification (CMC) methods perform well on moder…
Few-Shot Connectivity-Aware Text Line Segmentation in Historical Documents
Rafael Sterzinger, Tingyu Lin, Robert Sablatnig
A foundational task for the digital analysis of documents is text line segmentation. However, automating this process with deep learning models is challenging because it requires l…
Goku: Flow Based Video Generative Foundation Models
Shoufa Chen, Chongjian Ge, Yuqi Zhang +19
This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We…