8 papers
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani +5
Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, t…
ALARM: Audio-Language Alignment for Reasoning Models
Petr Grinberg, Hassan Shahmohammadi
Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fa…
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
Hayato Futami, Emiru Tsunoo, Yosuke Kashiwagi +4
Speech-to-speech translation (S2ST) has been advanced with large language models (LLMs), which are fine-tuned on discrete speech units. In such approaches, modality adaptation from…
ViPE: Visualise Pretty-much Everything
Hassan Shahmohammadi, Adhiraj Ghosh, Hendrik P. A. Lensch
Figurative and non-literal expressions are profoundly integrated in human communication. Visualising such expressions allow us to convey our creative thoughts, and evoke nuanced em…
Visual Grounding of Inter-lingual Word-Embeddings
Wafaa Mohammed, Hassan Shahmohammadi, Hendrik P. A. Lensch +1
Visual grounding of Language aims at enriching textual representations of language with multiple sources of visual knowledge such as images and videos. Although visual grounding is…
How direct is the link between words and images?
Hassan Shahmohammadi, Maria Heitmeier, Elnaz Shafaei-Bajestan +2
Current word embedding models despite their success, still suffer from their lack of grounding in the real world. In this line of research, Gunther et al. 2022 proposed a behaviora…