papers

Publications (12)

cs.AI2026

MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization

Yichen Han, Yuhang Han, Siteng Huang +7

Prompt engineering is crucial for fully leveraging large language models (LLMs), yet most existing optimization methods follow a single trajectory, resulting in limited adaptabilit…

cs.CV2022

A Keypoint Based Enhancement Method for Audio Driven Free View Talking Head Synthesis

Yichen Han, Ya Li, Yingming Gao +3

Audio driven talking head synthesis is a challenging task that attracts increasing attention in recent years. Although existing methods based on 2D landmarks or 3D face models can…

cs.AI2026

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

Bo Shen, Lifeng Chang, Tianyuan Wei +7

The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI thr…

cs.HC2023

Screen or No Screen? Lessons Learnt from a Real-World Deployment Study of Using Voice Assistants With and Without Touchscreen for Older Adults

Chen Chen, Ella T. Lifset, Yichen Han +5

While voice user interfaces offer increased accessibility due to hands-free and eyes-free interactions, older adults often have challenges such as constructing structured requests…

cs.CL2023

CONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis

Yayue Deng, Jinlong Xue, Yukang Jia +6

Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While p…

cs.HC2022

Towards Visualization of Time-Series Ecological Momentary Assessment (EMA) Data on Standalone Voice-First Virtual Assistants

Yichen Han, Christopher Bo Han, Chen Chen +5

Population aging is an increasingly important consideration for health care in the 21th century, and continuing to have access and interact with digital health information is a key…

cs.SD2025

MBCodec:Thorough disentangle for high-fidelity audio compression

Ruonan Zhang, Xiaoyang Hao, Yichen Han +3

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced…

cs.SD2025

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Yichen Han, Xiaoyang Hao, Keming Chen +25

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, whi…

cs.SD2022

ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis

Jinlong Xue, Yayue Deng, Yichen Han +3

In recent years, neural network based methods for multi-speaker text-to-speech synthesis (TTS) have made significant progress. However, the current speaker encoder models used in t…

cs.SD2025

Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling

Junjie Cao, Yichen Han, Ruonan Zhang +5

Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation…

cs.HC2024

How do Older Adults Set Up Voice Assistants? Lessons Learned from a Deployment Experience for Older Adults to Set Up Standalone Voice Assistants

Chen Chen, Ella T. Lifset, Yichen Han +5

While standalone Voice Assistants (VAs) are promising to support older adults' daily routine and wellbeing management, onboarding and setting up these devices can be challenging. A…

cs.SD2023

Frame-level emotional state alignment method for speech emotion recognition

Qifei Li, Yingming Gao, Cong Wang +4

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level la…