From the 1 of 9 linked papers with an AI index.
9 papers
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning
Zhiyue Zhao, Jingyi Wu, Hairuo Liu +5
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spac…
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Zhenqi Jia, Yuan Zhao, Aruukhan +2
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods str…
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
Yuan Zhao, Zhenqi Jia, Yongqiang Zhang
The paper introduces MAR3, a training‑free multi‑agent framework that uses large language model agents to recognize expression difficulty, determine dominant modality, reason about…
Emotion and Acoustics Should Agree: Cross-Level Inconsistency Analysis for Audio Deepfake Detection
Jinhua Zhang, Zhenqi Jia, Rui Liu
Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional featur…
Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning
Rui Liu, Yuan Zhao, Zhenqi Jia
The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video…
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
Zhenqi Jia, Rui Liu, Berrak Sisman +1
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate pro…