33 citations · 36 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 2 cited
Faster Diffusion via Temporal Attention Decomposition
Haozhe Liu, Wentian Zhang, Jinheng Xie +6
We explore the role of attention mechanism during inference in text-conditional diffusion models. Empirical observations suggest that cross-attention outputs converge to a fixed po…
cs.CV2023★ 1 cited
ViT-Lens: Towards Omni-modal Representations
Weixian Lei, Yixiao Ge, Kun Yi +6
Aiming to advance AI agents, large foundation models significantly improve reasoning and instruction execution, yet the current focus on vision and language neglects the potential…
cs.CV2021★ 33 cited
AVA-AVD: Audio-Visual Speaker Diarization in the Wild
Eric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui +3
Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor…