9 papers
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
Zijun Li, Yimin Zhou, Jia Sun +12
Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when…
RePer-360: Releasing Perspective Priors for 360 Depth Estimation via Self-Modulation
Cheng Guan, Chunyu Lin, Zhijie Shen +2
Recent depth foundation models trained on perspective imagery achieve strong performance, yet generalize poorly to 360 images due to the substantial geometric discrepancy b…
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Jiyuan Wang, Chunyu Lin, Lei Sun +8
Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains challenging in edited results, and the extr…
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin +15
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global tem…
From Editor to Dense Geometry Estimator
JiYuan Wang, Chunyu Lin, Lei Sun +5
Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image tas…
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
Jiyuan Wang, Chunyu Lin, Cheng Guan +5
In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to…