27 papers
Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation
Yasser Hamidullah, Koel Dutta Chowdhury, Yusser Al Ghussin +4
Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly critical in sign language transl…
The Latin Substrate: How Language Models Represent and Mediate Script Choice
Daniil Gurgurov, Alan Saji, Katharina Trinley +2
Many languages are written in multiple scripts, requiring large language models (LLMs) to generate equivalent linguistic content in distinct orthographic forms. While prior work su…
DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge
Yusser Al Ghussin, Daniil Gurgurov, Yasser Hamidullah +3
Large language models (LLMs) are increasingly used across diverse linguistic and cultural contexts, yet their cultural knowledge remains uneven across regions and languages. We pre…
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel +3
Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unrelia…
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding
Cennet Oguz, Yasser Hamidullah, Josef van Genabith +1
We introduce DualFact, a dual-layer, multimodal factuality evaluation framework for procedural video captioning. DualFact separates factual correctness into conceptual facts, captu…
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
Dan Shi, Zhuowen Han, Simon Ostermann +3
Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (S…