4 papers
Elastic Time: Dynamic Frame Rate Bottlenecks for Neural Audio Coding
Dimitrios Bralios, Paris Smaragdis, Minje Kim
Neural audio autoencoders have become a core component of compression, feature extraction, and generation. However, while existing systems support variable bitrate, the vast majori…
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim +6
Retrieval-augmented generation (RAG) has become a common practice in multimodal large language models (MLLM) to enhance factual grounding and reduce hallucination. Yet, its relianc…
Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders
Dimitrios Bralios, Jonah Casebeer, Paris Smaragdis
Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitatio…
Learning to Upsample and Upmix Audio in the Latent Domain
Dimitrios Bralios, Paris Smaragdis, Jonah Casebeer
Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression system…