3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2024
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
Yihan Wu, Soumi Maiti, Yifan Peng +6
Recent advancements in language models have significantly enhanced performance in multiple speech-related tasks. Existing speech language models typically utilize task-dependent pr…
cs.CV2023
ViCo: Engaging Video Comment Generation with Human Preference Rewards
Yuchong Sun, Bei Liu, Xu Chen +2
Engaging video comments play an important role in video social media, as they are the carrier of feelings, thoughts, or humor of the audience. Preliminary works have made initial e…
cs.CV2021★ 3 cited
Class-aware Sounding Objects Localization via Audiovisual Correspondence
Di Hu, Yake Wei, Rui Qian +3
Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achie…