5 papers
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
Youxin Pang, Jiajun Liu, Lingfeng Tan +6
We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limite…
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
Youxin Pang, Yong Zhang, Ruizhi Shao +5
We propose UniMo, an innovative autoregressive model for joint modeling of 2D human videos and 3D human motions within a unified framework, enabling simultaneous generation and und…
GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
Ken Deng, Yunhan Yang, Jingxiang Sun +4
We introduce GeoSAM2, a prompt-controllable framework for 3D part segmentation that casts the task as multi-view 2D mask prediction. Given a textureless object, we render normal an…
Interspatial Attention for Efficient 4D Human Video Generation
Ruizhi Shao, Yinghao Xu, Yujun Shen +5
Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ templ…
DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow
Ken Deng, Yuan-Chen Guo, Jingxiang Sun +6
Modern 3D generation methods can rapidly create shapes from sparse or single views, but their outputs often lack geometric detail due to computational constraints. We present Detai…