most citedLDM3D: Latent Diffusion Model for 3D

12 citations · 16 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024

Why do LLaVA Vision-Language Models Reply to Images in English?

Musashi Hinck, Carolin Holtermann, Matthew Lyle Olson +6

We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significant…

cs.CV20242 cited

L-MAGIC: Language Model Assisted Generation of Images with Coherence

Zhipeng Cai, Matthias Mueller, Reiner Birkl +6

In the current era of generative AI breakthroughs, generating panoramic scenes from a single input image remains a key challenge. Most existing methods use diffusion-based iterativ…

cs.CV20232 cited

LDM3D-VR: Latent Diffusion Model for 3D VR

Gabriela Ben Melech Stan, Diana Wofk, Estelle Aflalo +4

Latent diffusion models have proven to be state-of-the-art in the creation and manipulation of visual outputs. However, as far as we know, the generation of depth maps jointly with…

cs.CV2023

ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning

Xiao Xu, Bei Li, Chenfei Wu +6

Two-Tower Vision-Language (VL) models have shown promising improvements on various downstream VL tasks. Although the most advanced work improves performance by building bridges bet…

cs.CV202312 cited

LDM3D: Latent Diffusion Model for 3D

Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox +8

This research paper proposes a Latent Diffusion Model for 3D (LDM3D) that generates both image and depth map data from a given text prompt, allowing users to generate RGBD images f…