2 papers
cs.CV2024
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
Zhenggang Tang, Peiye Zhuang, Chaoyang Wang +5
The task of image-to-multi-view generation refers to generating novel views of an instance from a single image. Recent methods achieve this by extending text-to-image latent diffus…
cs.CV2024
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace +8
The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is mu…