11 citations · 19 across the 2 of their papers we have counts for
2 papers
cs.CV2017★ 8 cited
CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation
Wangli Hao, Zhaoxiang Zhang, He Guan
Visual and audio modalities are two symbiotic modalities underlying videos, which contain both common and complementary information. If they can be mined and fused sufficiently, pe…
cs.CV2017★ 11 cited
Integrating both Visual and Audio Cues for Enhanced Video Caption
Wangli Hao, Zhaoxiang Zhang, He Guan +1
Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existi…