95 citations · 136 across the 5 of their papers we have counts for
Showing 2023Show all
3 papers · 1 filter
cs.MM2023★ 1 cited
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
Jinzheng Zhao, Yong Xu, Xinyuan Qian +6
Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide compl…
cs.SD2023★ 34 cited
Multimodal Fish Feeding Intensity Assessment in Aquaculture
Meng Cui, Xubo Liu, Haohe Liu +5
Fish feeding intensity assessment (FFIA) aims to evaluate fish appetite changes during feeding, which is crucial in industrial aquaculture applications. Existing FFIA methods are l…
cs.SD2023★ 6 cited
WavJourney: Compositional Audio Creation with Large Language Models
Xubo Liu, Zhongkai Zhu, Haohe Liu +8
Despite breakthroughs in audio generation models, their capabilities are often confined to domain-specific conditions such as speech transcriptions and audio captions. However, rea…