collaborators

7 papers

cs.MA2026

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

Mingkuan Feng, Jinyang Wu, Hao Gu +5

Agentic multimodal models perform diverse operations on an image via code and reason over the returned view, an effective paradigm for fine-grained visual question answering. Howev…

cs.SD2026

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

Yong Ren, Jingbei Li, Haiyang Sun +6

Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence…

cs.SD2026

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

Yong Ren, Jiangyan Yi, Jianhua Tao +5

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a dir…

cs.SD2025

ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection

Hao Gu, Jiangyan Yi, Chenglong Wang +6

Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large lang…

cs.SD2025

Manipulated Regions Localization For Partially Deepfake Audio: A Survey

Jiayi He, Jiangyan Yi, Jianhua Tao +2

With the development of audio deepfake techniques, attacks with partially deepfake audio are beginning to rise. Compared to fully deepfake, it is much harder to be identified by th…

cs.MM2025

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu +7

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…