activity
20242026
most citedSAM 3D: 3Dfy Anything in Images

1 citations · 1 across the 1 of their papers we have counts for

collaborators

7 papers

cs.CV20261 cited

SAM 3D: 3Dfy Anything in Images

SAM 3D Team, Xingyu Chen, Fu-Jen Chu +20

We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images,…

cs.LG2026

Antithetic Noise in Diffusion Models

Jing Jia, Sifan Liu, Bowen Song +3

We systematically study antithetic initial noise in diffusion models, discovering that pairing each noise sample with its negation consistently produces strong negative correlation…

cs.CV2025

Learning Image Priors through Patch-based Diffusion Models for Solving Inverse Problems

Jason Hu, Bowen Song, Xiaojian Xu +2

Diffusion models can learn strong image priors from underlying data distribution and use them to solve inverse problems, but the training process is computationally expensive and r…

cs.LG2025

CCS: Controllable and Constrained Sampling with Diffusion Models via Initial Noise Perturbation

Bowen Song, Zecheng Zhang, Zhaoxu Luo +6

Diffusion models have emerged as powerful tools for generative tasks, producing high-quality outputs across diverse domains. However, how the generated data responds to the initial…

cs.CV2024

SatDiffMoE: A Mixture of Estimation Method for Satellite Image Super-resolution with Latent Diffusion Models

Zhaoxu Luo, Bowen Song, Liyue Shen

During the acquisition of satellite images, there is generally a trade-off between spatial resolution and temporal resolution (acquisition frequency) due to the onboard sensors of…

cs.CV2024

Latent Space Disentanglement in Diffusion Transformers Enables Precise Zero-shot Semantic Editing

Zitao Shuai, Chenwei Wu, Zhengxu Tang +2

Diffusion Transformers (DiTs) have recently achieved remarkable success in text-guided image generation. In image editing, DiTs project text and image inputs to a joint latent spac…