39 citations · 60 across the 30 of their papers we have counts for
Showing 2023 · cs.CVShow all
3 papers · 2 filters
cs.CV2023★ 8 cited
Video Understanding with Large Language Models: A Survey
Yolo Y. Tang, Jing Bi, Siting Xu +17
With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…
cs.CV2023★ 2 cited
High-Quality Visually-Guided Sound Separation from Diverse Categories
Chao Huang, Susan Liang, Yapeng Tian +2
We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typica…
cs.CV2023★ 6 cited
AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis
Susan Liang, Chao Huang, Yapeng Tian +2
Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task…