1 citations · 1 across the 9 of their papers we have counts for
5 papers · 1 filter
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
Jinkun Hao, Naifu Liang, Zhen Luo +8
The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, tradition…
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
Yuji Wang, Moran Li, Xiaobin Hu +7
Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to i…
Semantic Frame Interpolation
Yijia Hong, Jiangning Zhang, Ran Yi +4
Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significant research and application poten…
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
Yuji Wang, Moran Li, Xiaobin Hu +7
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. Howe…
3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations
Yating Wang, Xuan Wang, Ran Yi +4
Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture…