34 citations · 41 across the 7 of their papers we have counts for
12 papers
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Dibyadip Chatterjee, Edoardo Remelli, Yale Song +9
We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - ve…
SignRep: Enhancing Self-Supervised Sign Representations
Ryan Wong, Necati Cihan Camgoz, Richard Bowden
Sign language representation learning presents unique challenges due to the complex spatio-temporal nature of signs and the scarcity of labeled datasets. Existing methods often rel…
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
Marta R. Costa-jussà, Bokai Yu, Pierre Andrews +7
We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the inters…
Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation
Ryan Wong, Necati Cihan Camgoz, Richard Bowden
Automatic Sign Language Translation requires the integration of both computer vision and natural language processing to effectively bridge the communication gap between sign and sp…
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
Ozge Mercanoglu Sincan, Richard Bowden
Sign Language Translation (SLT) is a challenging task that aims to generate spoken language sentences from sign language videos. In this paper, we introduce a lightweight, modular…
Multi-channel Transformers for Multi-articulatory Sign Language Translation
Necati Cihan Camgoz, Oscar Koller, Simon Hadfield +1
Sign languages use multiple asynchronous information channels (articulators), not just the hands but also the face and body, which computational approaches often ignore. In this pa…