16 papers
DepthMaster: Taming Diffusion Models for Monocular Depth Estimation
Ziyang Song, Zerong Wang, Bo Li +5
Monocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference speed. Recent methods adopt a singl…
HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions
Shuolin Xu, Siming Zheng, Ziyi Wang +7
Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods…
Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution
Tianyi Zhang, Zheng-Peng Duan, Peng-Tao Jiang +4
Diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance.To achieve efficient Real-ISR, many works employ Variational Score Dis…
Photography Perspective Composition: Towards Aesthetic Perspective Recommendation
Lujian Yao, Siming Zheng, Xinbin Yuan +5
Traditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professio…
Learning Differential Pyramid Representation for Tone Mapping
Qirui Yang, Yinbo Li, Yihao Liu +5
Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. These designs typically fail to preserve fine textur…
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
Guangyuan Li, Siming Zheng, Hao Zhang +6
Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despit…