12 papers
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
Jiaqi Wang, Weijia Wu, Yi Zhan +6
With AI-generated videos increasingly indistinguishable from reality, current benchmarks primarily focus on broad semantic alignment and basic physical consistency, offering limite…
MIND: Benchmarking Memory Consistency and Action Control in World Models
Yixuan Ye, Xuanyu Lu, Yuxin Jiang +7
World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address th…
MatDecompSDF: High-Fidelity 3D Shape and PBR Material Decomposition from Multi-View Images
Chengyu Wang, Isabella Bennett, Henry Scott +4
We present MatDecompSDF, a novel framework for recovering high-fidelity 3D shapes and decomposing their physically-based material properties from multi-view images. The core challe…
Glance: Accelerating Diffusion Models with 1 Sample
Zhuobai Dong, Rui Zhao, Songjie Wu +5
Diffusion models have achieved remarkable success in image generation, yet their deployment remains constrained by the heavy computational cost and the need for numerous inference…
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
Ropeway Liu, Hangjie Yuan, Bo Dong +6
Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effec…
TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation
Enming Zhang, Zhengyu Li, Yanru Wu +5
Recent advances in Vision Transformers (ViTs) have significantly advanced semantic segmentation performance. However, their adaptation to new target domains remains challenged by d…