collaborators

6 papers

cs.CV2026

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

Nando Metzger, Prune Truong, Goutam Bhat +2

The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present Elastic3D, a controllable, direct end-to-end method for upgrading a…

cs.CV2026

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

Orest Kupyn, Goutam Bhat, Philipp Henzler +3

Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric representations suitable for downstream use. Current video diffusion mo…

cs.CV2026

Stitched Value Model for Diffusion Alignment

Hyojun Go, Hyungjin Chung, Prune Truong +8

For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challen…

cs.CV2026

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

Hyojun Go, Dominik Narnhofer, Goutam Bhat +3

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could…

cs.CV2025

M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion

Nina Shvetsova, Goutam Bhat, Prune Truong +2

We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reproj…

cs.CV2025

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Fan-Yun Sun, Weiyu Liu, Siyi Gu +6

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demon…