3 papers
cs.CV2025
DA: Depth Anything in Any Direction
Haodong Li, Wangguangdong Zheng, Jing He +5
Panorama has a full FoV (360180), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimati…
cs.CV2025
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
Jiyuan Wang, Chunyu Lin, Cheng Guan +5
In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to…
cs.CV2024
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
Leheng Li, Weichao Qiu, Xu Yan +6
We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be…