collaborators

6 papers

cs.MM2024

Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation

Haoxiang Shi, Ziqi Liang, Jun Yu

Emotion Prediction in Conversation (EPC) aims to forecast the emotions of forthcoming utterances by utilizing preceding dialogues. Previous EPC approaches relied on simple context…

cs.CL2024

QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question Answering

Sheng Ouyang, Jianzong Wang, Yong Zhang +5

Extractive Question Answering (EQA) in Machine Reading Comprehension (MRC) often faces the challenge of dealing with semantically identical but format-variant inputs. Our work intr…

cs.SD2024

EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization

Jianzong Wang, Ziqi Liang, Xulong Zhang +2

In recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storag…

cs.SD2024

EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning

Ziqi Liang, Jianzong Wang, Xulong Zhang +3

Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into a…

cs.SD2024

EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech

Ziqi Liang, Haoxiang Shi, Jiawei Wang +1

Recently, deep learning-based Text-to-Speech (TTS) systems have achieved high-quality speech synthesis results. Recurrent neural networks have become a standard modeling technique…

cs.CV2023

CP-EB: Talking Face Generation with Controllable Pose and Eye Blinking Embedding

Jianzong Wang, Yimin Deng, Ziqi Liang +3

This paper proposes a talking face generation method named "CP-EB" that takes an audio signal as input and a person image as reference, to synthesize a photo-realistic people talki…