2 papers
eess.AS2026
HuPER: A Human-Inspired Framework for Phonetic Perception
Chenxu Guo, Jiachen Lian, Yisi Liu +4
We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of…
cs.CV2025
Sounding that Object: Interactive Object-Aware Image to Audio Generation
Tingle Li, Baihe Huang, Xiaobin Zhuang +6
Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em inter…