From the 1 of 11 linked papers with an AI index.
11 papers
Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3
The paper introduces GMM-EVA, a training-free framework that uses Gaussian Mixture Models to detect event-level structures in long videos and allocate visual tokens by selecting on…
Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation
Siyi Chen, Shaowei Liu, Yixuan Jia +4
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its suc…
PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement
Jun Gao, Xiaobin Rong, Yu Sun +2
Flow matching (FM) enables high-fidelity generation, while self-supervised learning (SSL) speech models provide hierarchical representations spanning acoustic and phonetic levels.…
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
Xiaobin Rong, Zheng Wang, Yushi Wang +2
Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination…
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
Dahan Wang, Jun Gao, Tong Lei +4
Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech si…
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
Chaojie Yang, Tian Li, Yue Zhang +1
Diffusion Transformer (DiT) architectures have significantly advanced Text-to-Image (T2I) generation but suffer from prohibitive computational costs and deployment barriers. To add…