5 papers
Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization
Yudong Li, Zihao Fang, Junwen Qiu +4
Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bot…
A Survey of Large Language Models for Perception and Measurement of Human Psychology
Yudong Li, Xiaoyi Chen, Jiawei Cai +4
Against the backdrop of the rapid advancement of Large Language Models (LLMs), their application in the field of psychology has garnered significant academic attention. A central i…
KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
Yudong Li, Jiawei Cai, Linlin Shen
Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that humans naturally rely on to co…
GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
Yudong Li, Hao Li, Xianxu Hou +1
Compared to the prosperity of pre-training models in natural image understanding, the research on large-scale pre-training models for facial knowledge learning is still limited. Cu…
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
Xiaoqin Wang, Xusen Ma, Xianxu Hou +6
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely…