21 citations · 37 across the 15 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.CV2022★ 3 cited
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
Xiao Xu, Chenfei Wu, Shachar Rosenman +3
Vision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years. Current VL models either use lightweight uni-mod…
cs.CL2022
Text is no more Enough! A Benchmark for Profile-based Spoken Language Understanding
Xiao Xu, Libo Qin, Kaiji Chen +3
Current researches on spoken language understanding (SLU) heavily are limited to a simple setting: the plain text-based SLU that takes the user utterance as input and generates its…