13 citations · 40 across the 22 of their papers we have counts for
6 papers · 1 filter
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
Xinwang Chen, Ning Liu, Yichen Zhu +2
Transformer-based Diffusion Probabilistic Models (DPMs) have shown more potential than CNN-based DPMs, yet their extensive computational requirements hinder widespread practical ap…
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
Minjie Zhu, Yichen Zhu, Xin Liu +7
Multimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles…
DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template Field
Haowen Wang, Zhipeng Fan, Zhen Zhao +7
Estimating 6D poses and reconstructing 3D shapes of objects in open-world scenes from RGB-depth image pairs is challenging. Many existing methods rely on learning geometric feature…
RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion
Haowen Wang, Zhengping Che, Yufan Yang +6
Raw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent…
CP: Channel Pruning Plug-in for Point-based Networks
Yaomin Huang, Ning Liu, Zhengping Che +7
Channel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success ha…
RGB-Depth Fusion GAN for Indoor Depth Completion
Haowen Wang, Mingyuan Wang, Zhengping Che +5
The raw depth image captured by the indoor depth sensor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transpa…