most citedEnhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning

Xuyang Zhao, Liting Zhang, Zichen Xu +4

Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. H…

cs.SD2026

EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry

Zhongyuan Fu, Yuhang Jia, Hui Wang +6

Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits re…

cs.SD2024

AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework

Yuhang Jia, Yang Chen, Jinghua Zhao +4

Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audi…

cs.SD2024

M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper

Jiaming Zhou, Shiwan Zhao, Jiabei He +6

State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizin…

cs.MM20241 cited

Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework

Haoqin Sun, Shiwan Zhao, Shaokai Li +7

Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this…