5 papers
Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs
Andreas Maier, Jeta Sopa, Gozde Gul Sahin +2
Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system prompt contains a soft sponsorshi…
Safety and accuracy follow different scaling laws in clinical large language models
Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa +9
Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation that higher accuracy implies…
Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
Mina Farajiamiri, Jeta Sopa, Saba Afza +9
Agentic retrieval-augmented reasoning pipelines are increasingly used to structure how large language models (LLMs) incorporate external evidence in clinical decision support. Thes…
SteuerLLM: Local specialized large language model for German tax law analysis
Sebastian Wind, Jeta Sopa, Laurin Schmid +8
Large language models (LLMs) demonstrate strong general reasoning and language understanding, yet their performance degrades in domains governed by strict formal rules, precise ter…
Multi-step retrieval and reasoning improves radiology question answering with large language models
Sebastian Wind, Jeta Sopa, Daniel Truhn +9
Clinical decision-making in radiology increasingly benefits from artificial intelligence (AI), particularly through large language models (LLMs). However, traditional retrieval-aug…