8 citations · 16 across the 13 of their papers we have counts for
2 papers
cs.CV2022★ 8 cited
CLIP-Event: Connecting Text and Images with Event Structures
Manling Li, Ruochen Xu, Shuohang Wang +6
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing v…
cs.LG2021
Sketching as a Tool for Understanding and Accelerating Self-attention for Long Sequences
Yifan Chen, Qi Zeng, Dilek Hakkani-Tur +3
Transformer-based models are not efficient in processing long sequences due to the quadratic space and time complexity of the self-attention modules. To address this limitation, Li…