3 papers
cs.AI2026
The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation
Doan Nam Long Vu, Simone Balloccu
Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evaluate 12 open-weight vision-language models…
cs.HC2026
Subjective Code Preferences in Experts and Large Language Models
Anna Mokhova, Subhabrata Dutta, Iryna Gurevych +1
Large Language Models (LLMs) have become increasingly popular for coding tasks, with subjective coding preferences being an essential element to adapt to programmers' personal need…
cs.CL2026
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
ZdenÄk Kasner, Vilém Zouhar, PatrÃcia Schmidtová +7
Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently…