25 citations · 35 across the 10 of their papers we have counts for
5 papers · 1 filter
GCT: Gated Contextual Transformer for Sequential Audio Tagging
Yuanbo Hou, Yun Wang, Wenwu Wang +1
Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class informa…
Relation-guided acoustic scene classification aided with event embeddings
Yuanbo Hou, Bo Kang, Wout Van Hauwermeiren +1
In real life, acoustic scenes and audio events are naturally correlated. Humans instinctively rely on fine-grained audio events as well as the overall sound characteristics to dist…
CT-SAT: Contextual Transformer for Sequential Audio Tagging
Yuanbo Hou, Zhaoyi Liu, Bo Kang +2
Sequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an au…
Attention-based cross-modal fusion for audio-visual voice activity detection in musical video streams
Yuanbo Hou, Zhesong Yu, Xia Liang +4
Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse mu…
Rule-embedded network for audio-visual voice activity detection in live musical video streams
Yuanbo Hou, Yi Deng, Bilei Zhu +2
Detecting anchor's voice in live musical streams is an important preprocessing for music and speech signal processing. Existing approaches to voice activity detection (VAD) primari…