Publications (12)
MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization
Yichen Han, Yuhang Han, Siteng Huang +7
Prompt engineering is crucial for fully leveraging large language models (LLMs), yet most existing optimization methods follow a single trajectory, resulting in limited adaptabilit…
A Keypoint Based Enhancement Method for Audio Driven Free View Talking Head Synthesis
Yichen Han, Ya Li, Yingming Gao +3
Audio driven talking head synthesis is a challenging task that attracts increasing attention in recent years. Although existing methods based on 2D landmarks or 3D face models can…
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
Bo Shen, Lifeng Chang, Tianyuan Wei +7
The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI thr…
Screen or No Screen? Lessons Learnt from a Real-World Deployment Study of Using Voice Assistants With and Without Touchscreen for Older Adults
Chen Chen, Ella T. Lifset, Yichen Han +5
While voice user interfaces offer increased accessibility due to hands-free and eyes-free interactions, older adults often have challenges such as constructing structured requests…
CONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis
Yayue Deng, Jinlong Xue, Yukang Jia +6
Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While p…
Towards Visualization of Time-Series Ecological Momentary Assessment (EMA) Data on Standalone Voice-First Virtual Assistants
Yichen Han, Christopher Bo Han, Chen Chen +5
Population aging is an increasingly important consideration for health care in the 21th century, and continuing to have access and interact with digital health information is a key…
MBCodec:Thorough disentangle for high-fidelity audio compression
Ruonan Zhang, Xiaoyang Hao, Yichen Han +3
High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced…
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
Yichen Han, Xiaoyang Hao, Keming Chen +25
Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, whi…
ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis
Jinlong Xue, Yayue Deng, Yichen Han +3
In recent years, neural network based methods for multi-speaker text-to-speech synthesis (TTS) have made significant progress. However, the current speaker encoder models used in t…
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
Junjie Cao, Yichen Han, Ruonan Zhang +5
Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation…
How do Older Adults Set Up Voice Assistants? Lessons Learned from a Deployment Experience for Older Adults to Set Up Standalone Voice Assistants
Chen Chen, Ella T. Lifset, Yichen Han +5
While standalone Voice Assistants (VAs) are promising to support older adults' daily routine and wellbeing management, onboarding and setting up these devices can be challenging. A…
Frame-level emotional state alignment method for speech emotion recognition
Qifei Li, Yingming Gao, Cong Wang +4
Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level la…