5 papers
Learning What to Attend First: Modality-Importance-Guided Reasoning for Reliable Multimodal Emotion Understanding
Hyeongseop Rha, Jeong Hun Yeo, Junil Won +2
In this paper, we present Modality-Importance-Guided Reasoning (MIGR), a framework designed to improve the reliability of reasoning-based multimodal emotion understanding in multim…
Long-Form Speech Generation with Spoken Language Models
Se Jin Park, Julian Salazar, Aren Jansen +3
We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken l…
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park +1
Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (…
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
Yeonju Kim, Se Jin Park, Yong Man Ro
Chatbot research is advancing with the growing importance of chatbots in fields that require human interactions, such as customer support and mental health care. Despite these adva…
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
Se Jin Park, Yeonju Kim, Hyeongseop Rha +2
In human communication, both verbal and non-verbal cues play a crucial role in conveying emotions, intentions, and meaning beyond words alone. These non-linguistic information, suc…