activity
20222024
most citedDDMM-Synth: A Denoising Diffusion Model for Cross-modal Medical Image Synthesis with Sparse-view Measurement Embedding

9 citations · 35 across the 23 of their papers we have counts for

collaborators

12 papers

cs.CV2023

Vision meets mmWave Radar: 3D Object Perception Benchmark for Autonomous Driving

Yizhou Wang, Jen-Hao Cheng, Jui-Te Huang +8

Sensor fusion is crucial for an accurate and robust perception system on autonomous vehicles. Most existing datasets and perception solutions focus on fusing cameras and LiDAR. How…

cs.LG2023

Devil in the Number: Towards Robust Multi-modality Data Filter

Yichen Xu, Zihan Xu, Wenhao Chai +3

In order to appropriately filter multi-modality data sets on a web-scale, it becomes crucial to employ suitable filtering methods to boost performance and reduce training costs. Fo…

cs.CV2023

FrameRS: A Video Frame Compression Model Composed by Self supervised Video Frame Reconstructor and Key Frame Selector

Qiqian Fu, Guanhong Wang, Gaoang Wang

In this paper, we present frame reconstruction model: FrameRS. It consists self-supervised video frame reconstructor and key frame selector. The frame reconstructor, FrameMAE, is d…

cs.CV20232 cited

Chasing Consistency in Text-to-3D Generation from a Single Image

Yichen Ouyang, Wenhao Chai, Jiayi Ye +3

Text-to-3D generation from a single-view image is a popular but challenging task in 3D vision. Although numerous methods have been proposed, existing works still suffer from the in…

cs.CV2023

UniAP: Towards Universal Animal Perception in Vision via Few-shot Learning

Meiqi Sun, Zhonghan Zhao, Wenhao Chai +5

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is…

cs.CV20235 cited

StableVideo: Text-driven Consistency-aware Diffusion Video Editing

Wenhao Chai, Xun Guo, Gaoang Wang +1

Diffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents d…