1 paper
Anthony Chen, Naomi Ken Korem, Gal Zeevi +6
Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and…