3 papers
cs.CV2024
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
Liuhan Chen, Zongjian Li, Bin Lin +6
Variational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). With the same reconstructi…
cs.CV2024
Envision3D: One Image to 3D with Anchor Views Interpolation
Yatian Pang, Tanghui Jia, Yujun Shi +6
We present Envision3D, a novel method for efficiently generating high-quality 3D content from a single image. Recent methods that extract 3D content from multi-view images generate…
cs.CL2024
LLMBind: A Unified Modality-Task Integration Framework
Bin Zhu, Munan Ning, Peng Jin +7
Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generat…