activity
20242026
collaborators

5 papers

cs.LG2026

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos +2

Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability wo…

cs.CL2025

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos +1

We present a multimodal fusion framework that bridges pre-trained decoder-based large language models (LLM) and acoustic encoder-decoder architectures such as Whisper, with the aim…

cs.CL2025

Krikri: Advancing Open Large Language Models for Greek

Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos +6

We introduce Llama-Krikri-8B, a cutting-edge Large Language Model tailored for the Greek language, built on Meta's Llama 3.1-8B. Llama-Krikri-8B has been extensively trained on hig…

cs.CL2024

Meltemi: The first open Large Language Model for Greek

Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos +6

We describe the development and capabilities of Meltemi 7B, the first open Large Language Model for the Greek language. Meltemi 7B has 7 billion parameters and is trained on a 40 b…

cs.CL2024

The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data

Georgios Paraskevopoulos, Chara Tsoukala, Athanasios Katsamanis +1

The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue i…