4 papers
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
Griffin Farrow, Lily Sijia Li, Jack Johnson +4
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading…
CoRet: Improved Retriever for Code Editing
Fabio Fehr, Prabhu Teja Sivaprasad, Luca Franceschi +1
In this paper, we introduce CoRet, a dense retrieval model designed for code-editing tasks that integrates code semantics, repository structure, and call graph dependencies. The mo…
Nonparametric Variational Regularisation of Pretrained Transformers
Fabio Fehr, James Henderson
The current paradigm of large-scale pre-training and fine-tuning Transformer large language models has lead to significant improvements across the board in natural language process…
Learning to Abstract with Nonparametric Variational Information Bottleneck
Melika Behjati, Fabio Fehr, James Henderson
Learned representations at the level of characters, sub-words, words and sentences, have each contributed to advances in understanding different NLP tasks and linguistic phenomena.…