5 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 2 cited
Linearized Relative Positional Encoding
Zhen Qin, Weixuan Sun, Kaiyue Lu +6
Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are…
cs.CL2023★ 5 cited
Toeplitz Neural Network for Sequence Modeling
Zhen Qin, Xiaodong Han, Weixuan Sun +6
Sequence modeling has important applications in natural language processing and computer vision. Recently, the transformer-based models have shown strong performance on various seq…
cs.CV2023
Fine-grained Audible Video Description
Xuyang Shen, Dong Li, Jinxing Zhou +9
We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audibl…