5 papers
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos +2
Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability wo…
VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos +1
We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim…
Krikri: Advancing Open Large Language Models for Greek
Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos +6
We introduce Llama-Krikri-8B, a cutting-edge Large Language Model tailored for the Greek language, built on Meta's Llama 3.1-8B. Llama-Krikri-8B has been extensively trained on hig…
Meltemi: The first open Large Language Model for Greek
Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos +6
We describe the development and capabilities of Meltemi 7B, the first open Large Language Model for the Greek language. Meltemi 7B has 7 billion parameters and is trained on a 40 b…
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
Georgios Paraskevopoulos, Chara Tsoukala, Athanasios Katsamanis +1
The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue i…