2 papers
cs.CV2025
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
Zeteng Lin, Xingxing Li, Wen You +4
Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this…
cs.SD2025
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
Advait Gosai, Tyler Vuong, Utkarsh Tyagi +8
End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate tran…