1 citations · 1 across the 7 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
Junbo Wang, Haofeng Tan, Bowen Liao +7
Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to rec…
cs.SD2025
Voxtral
Alexander H. Liu, Andy Ehrenberg, Andy Lo +103
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art perfo…