collaborators

16 papers

cs.CL2025

Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs

Juraj Vladika, Annika Domres, Mai Nguyen +10

Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and inaccurate citations, which pose a challenge to their clinical adoption and reg…

cs.CL2025

Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study

Mahdi Dhaini, Juraj Vladika, Ege Erdogan +2

In the rapidly evolving field of Explainable Natural Language Processing (NLP), textual explanations, i.e., human-like rationales, are pivotal for explaining model predictions and…

cs.CL2025

Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models

Juraj Vladika, Mahdi Dhaini, Florian Matthes

The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance o…

cs.CL2025

FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain

Anum Afzal, Juraj Vladika, Florian Matthes

Large Language Models tend to struggle when dealing with specialized domains. While all aspects of evaluation hold importance, factuality is the most critical one. Similarly, relia…

cs.CL2025

MedSEBA: Synthesizing Evidence-Based Answers Grounded in Evolving Medical Literature

Juraj Vladika, Florian Matthes

In the digital age, people often turn to the Internet in search of medical advice and recommendations. With the increasing volume of online content, it has become difficult to dist…

cs.CL2025

Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge

Juraj Vladika, Ihsan Soydemir, Florian Matthes

While large language models (LLMs) have shown remarkable capabilities to generate coherent text, they suffer from the issue of hallucinations -- factually inaccurate statements. Am…