4 papers
Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG
Ilias Triantafyllopoulos, Renyi Qu, Salvatore Giorgi +3
Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains, where safety depends not only on how a system answers, but also on whether a query sh…
Conceptors for Semantic Steering
Ilias Triantafyllopoulos, Young-Min Cho, Ren Tao +6
Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction whose geometry is left largely un…
Learning to Pay Attention: Unsupervised Modeling of Attentive and Inattentive Respondents in Survey Data
Ilias Triantafyllopoulos, Panos Ipeirotis
The integrity of behavioral and social-science surveys depends on detecting inattentive respondents who provide random or low-effort answers. Traditional safeguards, such as attent…
Interpreting and Mitigating Unwanted Uncertainty in LLMs
Tiasa Singha Roy, Ayush Rajesh Jhaveri, Ilias Triantafyllopoulos
Despite their impressive capabilities, Large Language Models (LLMs) exhibit unwanted uncertainty, a phenomenon where a model changes a previously correct answer into an incorrect o…