4 papers
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
Xuanchen Li, Chenrui Cui, Tianrui Wang +9
Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across sourc…
SingingBot: An Avatar-Driven System for Robotic Face Singing Performance
Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng +3
Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversati…
MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image
DongFu Yin, Xiaotian Chen, Fei Richard Yu +2
Advances in generative modeling have significantly enhanced digital content creation, extending from 2D images to complex 3D and 4D scenes. Despite substantial progress, producing…
Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
Xuanchen Li, Jianyu Wang, Yuhao Cheng +5
Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In thi…