4 papers
4D Reconstruction from Sparse Dynamic Cameras
Kazuki Ozeki, Shun Kenney, Yuto Shibata +6
Although dynamic 3D (i.e., 4D) reconstruction from a monocular dynamic camera has recently advanced, it remains fundamentally limited by depth ambiguity. In this paper, we focus on…
MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
Akira Takahashi, Ryosuke Sawata, Shusuke Takahashi +1
Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as re…
Diffusion-based Signal Refiner for Speech Enhancement and Separation
Masato Hirano, Ryosuke Sawata, Naoki Murata +2
Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper propos…
HumanGif: Single-View Human Diffusion with Generative Prior
Shoukang Hu, Takuya Narihira, Kazumi Fukuda +3
Previous 3D human creation methods have made significant progress in synthesizing view-consistent and temporally aligned results from sparse-view images or monocular videos. Howeve…