55 citations · 83 across the 13 of their papers we have counts for
1 paper · 1 filter
Xinsheng Wang, Siyuan Feng, Jihua Zhu +2
This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the…