4 papers
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Le Xu, Chenxing Li, Yong Ren +5
Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge t…
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
Yong Ren, Chenxing Li, Manjie Xu +4
Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmon…
Video-to-Audio Generation with Hidden Alignment
Manjie Xu, Chenxing Li, Xinyi Tu +5
Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthr…
Improving Accuracy and Generalization for Efficient Visual Tracking
Ram Zaveri, Shivang Patel, Yu Gu +1
Efficient visual trackers overfit to their training distributions and lack generalization abilities, resulting in them performing well on their respective in-distribution (ID) test…