2 citations · 4 across the 12 of their papers we have counts for
3 papers · 1 filter
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
Burcu Sayin, Ngoc Vo Hong, Ipek Baris Schlicht +8
Clinical decision-making in emergency medicine demands rapid, accurate diagnoses under uncertainty. Despite benchmark progress, evidence for LLMs as interactive aids in live physic…
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
Sebastiano Monti, Carlo Nicolini, Gianni Pellegrini +2
Although the capabilities of large language models have been increasingly tested on complex reasoning tasks, their long-horizon planning abilities have not yet been extensively inv…
The LLM Wears Prada: Analysing Gender Bias and Stereotypes through Online Shopping Data
Massimiliano Luca, Ciro Beneduce, Bruno Lepri +1
With the wide and cross-domain adoption of Large Language Models, it becomes crucial to assess to which extent the statistical correlations in training data, which underlie their i…