4 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization
Edward Fish, Jon Weinbren, Andrew Gilbert
Temporal Action Localization (TAL) aims to identify actions' start, end, and class labels in untrimmed videos. While recent advancements using transformer networks and Feature Pyra…
cs.CV2022★ 4 cited
Two-Stream Transformer Architecture for Long Video Understanding
Edward Fish, Jon Weinbren, Andrew Gilbert
Pure vision transformer architectures are highly effective for short video classification and action recognition tasks. However, due to the quadratic complexity of self attention a…
cs.CV2021★ 3 cited
Rethinking movie genre classification with fine-grained semantic clustering
Edward Fish, Jon Weinbren, Andrew Gilbert
Movie genre classification is an active research area in machine learning. However, due to the limited labels available, there can be large semantic variations between movies withi…