11 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.SD2021★ 1 cited
End-to-end speaker diarization with transformer
Yongquan Lai, Xin Tang, Yuanyuan Fu +1
Speaker diarization is connected to semantic segmentation in computer vision. Inspired from MaskFormer \cite{cheng2021per} which treats semantic segmentation as a set-prediction pr…
cs.CV2021★ 11 cited
Visual-Semantic Transformer for Scene Text Recognition
Xin Tang, Yongquan Lai, Ying Liu +2
Modeling semantic information is helpful for scene text recognition. In this work, we propose to model semantic and visual information jointly with a Visual-Semantic Transformer (V…