most citedBoosting Fast and High-Quality Speech Synthesis with Linear Diffusion

1 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20241 cited

InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

Xiaotian Han, Yiren Jian, Xuefeng Hu +8

Pre-training on large-scale, high-quality datasets is crucial for enhancing the reasoning capabilities of Large Language Models (LLMs), especially in specialized domains such as ma…

cs.CV2024

InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Haogeng Liu, Quanzeng You, Xiaotian Han +7

Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intri…

cs.CV20231 cited

Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Haogeng Liu, Qihang Fan, Tingkai Liu +5

This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…

cs.SD20231 cited

Boosting Fast and High-Quality Speech Synthesis with Linear Diffusion

Haogeng Liu, Tao Wang, Jie Cao +2

Denoising Diffusion Probabilistic Models have shown extraordinary ability on various generative tasks. However, their slow inference speed renders them impractical in speech synthe…

cs.SD20231 cited

UnifySpeech: A Unified Framework for Zero-shot Text-to-Speech and Voice Conversion

Haogeng Liu, Tao Wang, Ruibo Fu +3

Text-to-speech (TTS) and voice conversion (VC) are two different tasks both aiming at generating high quality speaking voice according to different input modality. Due to their sim…