3 papers
cs.CV2025
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…
cs.CL2025
SSR: Alignment-Aware Modality Connector for Speech Language Models
Weiting Tan, Hirofumi Inaguma, Ning Dong +2
Fusing speech into pre-trained language model (SpeechLM) usually suffers from inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality. We…
eess.AS2025
Language model integration based on memory control for sequence to sequence speech recognition
Jaejin Cho, Shinji Watanabe, Takaaki Hori +4
In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained LM. Our proposed fusion methods focus on the memory cell state and the hidden stat…