2 papers
cs.SD2026
Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling
Anastasia Zorkina, Alexandr Anikin, Nikita Khmelev +5
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of…
eess.AS2025
Cryfish: On deep audio analysis with Large Language Models
Anton Mitrofanov, Sergei Novoselov, Tatiana Prisyach +8
The recent revolutionary progress in text-based large language models (LLMs) has contributed to the growth of interest in extending capabilities of such models to multimodal percep…