1 citations · 1 across the 5 of their papers we have counts for
5 papers
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
Luca Modica, Filip Landin, Mehrdad Farahani +3
In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechan…
Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning
Livia Qian, Gabriel Skantze
Backchannels (e.g., `yeah', `mhm', and `right') are short, non-interruptive feedback signals whose lexical form and prosody jointly convey pragmatic meaning. While prior computatio…
Representation of perceived prosodic similarity of conversational feedback
Livia Qian, Carol Figueroa, Gabriel Skantze
Vocal feedback (e.g., `mhm', `yeah', `okay') is an important component of spoken dialogue and is crucial to ensuring common ground in conversational systems. The exact meaning of s…
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
Livia Qian, Gabriel Skantze
Short feedback responses, such as backchannels, play an important role in spoken dialogue. So far, most of the modeling of feedback responses has focused on their timing, often neg…
Resolving References in Visually-Grounded Dialogue via Text Generation
Bram Willemsen, Livia Qian, Gabriel Skantze
Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge…