56 citations · 80 across the 2 of their papers we have counts for
2 papers
cs.SD2021★ 24 cited
Audio Description from Image by Modal Translation Network
Hailong Ning, Xiangtao Zheng, Yuan Yuan +1
Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to…
cs.MM2021★ 56 cited
Semantics-Consistent Representation Learning for Remote Sensing Image-Voice Retrieval
Hailong Ning, Bin Zhao, Yuan Yuan
With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image…