6 papers · 1 filter
What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces
Birger Moëll
What makes writing "good" remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evalua…
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
Birger Moell, Fredrik Sand Aronsson, Sanian Akbar
Integrating large language models (LLMs) like DeepSeek R1 into healthcare requires rigorous evaluation of their reasoning alignment with clinical expertise. This study assesses Dee…
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
Birger Moell, Fredrik Sand Aronsson, Per Ãstberg +1
Background:Speech patterns have emerged as potential diagnostic markers for conditions with varying etiologies. Machine learning (ML) presents an opportunity to harness these patte…
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
Birger Moell, Johan Boye
Large Language Models (LLMs) have made significant strides in natural language generation but often face challenges in tasks requiring precise calculations and structural analysis.…
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark
Birger Moell
In the rapidly evolving field of artificial intelligence, large language models (LLMs) have demonstrated significant capabilities across numerous applications. However, the perform…
Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support
Birger Moell
Background: Rapid advancements in natural language processing have led to the development of large language models with the potential to revolutionize mental health care. These mod…