Showing 2025 · cs.LGShow all
2 papers · 2 filters
cs.LG2025
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
Leon Eshuijs, Shihan Wang, Antske Fokkens
Reliance on spurious correlations (shortcuts) has been shown to underlie many of the successes of language models. Previous work focused on identifying the input elements that impa…
cs.LG2025
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
Leon Eshuijs, Archie Chaudhury, Alan McBeth +1
LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycop…