3 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 3 cited
Selective Structured State-Spaces for Long-Form Video Understanding
Jue Wang, Wentao Zhu, Pichao Wang +4
Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its lin…
cs.SD2023★ 1 cited
Multiscale Audio Spectrogram Transformer for Efficient Audio Classification
Wentao Zhu, Mohamed Omar
Audio event has a hierarchical architecture in both time and frequency and can be grouped together to construct more abstract semantic audio classes. In this work, we develop a mul…
cs.CV2022★ 3 cited
An End-to-End OCR Framework for Robust Arabic-Handwriting Recognition using a Novel Transformers-based Model and an Innovative 270 Million-Words Multi-Font Corpus of Classical Arabic with Diacritics
Aly Mostafa, Omar Mohamed, Ali Ashraf +4
This research is the second phase in a series of investigations on developing an Optical Character Recognition (OCR) of Arabic historical documents and examining how different mode…