collaborators

5 papers

cs.SD2026

FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations

Feiyu Shen, Kun Xie, Yichen Wu +8

Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capab…

cs.CV2026

PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On

Haohua Chen, Tianze Zhou, Wei Zhu +8

Virtual Try-on (VTON) has become a core capability for online retail, where realistic try-on results provide reliable fit guidance, reduce returns, and benefit both consumers and m…

eess.AS2026

FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System

Kaituo Xu, Yan Jia, Kai Huang +6

We present FireRedASR2S, a state-of-the-art industrial-grade all-in-one automatic speech recognition (ASR) system. It integrates four modules in a unified pipeline: ASR, Voice Acti…

cs.SD2025

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Junjie Chen, Yao Hu, Junjie Li +12

Full-duplex voice interaction allows users and agents to speak simultaneously with controllable barge-in, enabling lifelike assistants and customer service. Existing solutions are…

cs.SD2025

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Kun Xie, Feiyu Shen, Junjie Li +3

Current dialogue generation approaches typically require the complete dialogue text before synthesis and produce a single, inseparable speech containing all voices, making them uns…