1 paper
Wei-Ning Hsu, David Harwath, Christopher Song +1
In this paper we present the first model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an inter…