1 citations · 1 across the 3 of their papers we have counts for
9 papers
Glance: Accelerating Diffusion Models with 1 Sample
Zhuobai Dong, Rui Zhao, Songjie Wu +5
Diffusion models have achieved remarkable success in image generation, yet their deployment remains constrained by the heavy computational cost and the need for numerous inference…
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
Ropeway Liu, Hangjie Yuan, Bo Dong +6
Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effec…
MatDecompSDF: High-Fidelity 3D Shape and PBR Material Decomposition from Multi-View Images
Chengyu Wang, Isabella Bennett, Henry Scott +4
We present MatDecompSDF, a novel framework for recovering high-fidelity 3D shapes and decomposing their physically-based material properties from multi-view images. The core challe…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation
Enming Zhang, Zhengyu Li, Yanru Wu +5
Recent advances in Vision Transformers (ViTs) have significantly advanced semantic segmentation performance. However, their adaptation to new target domains remains challenged by d…
Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors
Yuke Lou, Yiming Wang, Zhen Wu +4
Human-object interaction (HOI) synthesis is important for various applications, ranging from virtual reality to robotics. However, acquiring 3D HOI data is challenging due to its c…