12 citations · 22 across the 4 of their papers we have counts for
4 papers
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
Zeyu Jin, Xiaoyu Qin, Songtao Zhou +2
Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an…
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
Qixin Wang, Songtao Zhou, Zeyu Jin +3
Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook es…
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
Yixuan Zhou, Xiaoyu Qin, Zeyu Jin +5
Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to spe…
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
Zeyu Jin, Jia Jia, Qixin Wang +5
Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elab…