4 papers
AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
Cheng Zhu, Jing Han, Qianshuai Xue +3
Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient f…
PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
Bin Wang, Yang Xu, Huan Zhao +2
Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high…
Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
Zhongren Dong, Haotian Guo, Weixiang Xu +2
Neuropsychiatric disorders, such as Alzheimer's disease (AD), depression, and autism spectrum disorder (ASD), are characterized by linguistic and acoustic abnormalities, offering p…
Celler:A Genomic Language Model for Long-Tailed Single-Cell Annotation
Huan Zhao, Yiming Liu, Jina Yao +3
Recent breakthroughs in single-cell technology have ushered in unparalleled opportunities to decode the molecular intricacy of intricate biological systems, especially those linked…