2 papers
cs.SD2026
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
Junbo Wang, Haofeng Tan, Bowen Liao +7
Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to rec…
cs.MM2025
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
Junbo Wang, Yan Zhao, Shuo Li +3
Micro-expressions (MEs) are crucial leakages of concealed emotion, yet their study has been constrained by a reliance on silent, visual-only data. To solve this issue, we introduce…