8 citations · 16 across the 48 of their papers we have counts for
6 papers · 1 filter
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding
Cennet Oguz, Yasser Hamidullah, Josef van Genabith +1
We introduce DualFact, a dual-layer, multimodal factuality evaluation framework for procedural video captioning. DualFact separates factual correctness into conceptual facts, captu…
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
Dan Shi, Zhuowen Han, Simon Ostermann +3
Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (S…
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
Tanja Baeumel, Josef van Genabith, Simon Ostermann
Large language models (LLMs) have demonstrated impressive capabilities, yet their internal mechanisms for handling reasoning-intensive tasks remain underexplored. To advance the un…
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
Simon Ostermann, Daniil Gurgurov, Tanja Baeumel +4
Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficient adaptation, and prompting…
ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
Daniil Gurgurov, Tom Röhr, Sebastian von Rohrscheidt +3
Despite advances in multilingual capabilities, most large language models (LLMs) remain English-centric in their training and, crucially, in their production of reasoning traces. E…
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel +5
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipula…