2 papers
cs.SD2025
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Ali Vosoughi, Yongyi Zang, Qihui Yang +3
Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: t…
cs.CV2025
Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
Jing Bi, Guangyu Sun, Ali Vosoughi +2
Multimodal large language models (MLLMs) that integrate visual and textual reasoning leverage chain-of-thought (CoT) prompting to tackle complex visual tasks, yet continue to exhib…