4 papers · 1 filter
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
Esma Balkır, Alice Pernthaller, Marco Basaldella +2
Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tas…
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
Sanhanat Sivapiromrat, Caiqi Zhang, Marco Basaldella +1
Recent studies have shown that Large Language Models (LLMs) are vulnerable to data poisoning attacks, where malicious training examples embed hidden behaviours triggered by specifi…
LUQ: Long-text Uncertainty Quantification for LLMs
Caiqi Zhang, Fangyu Liu, Marco Basaldella +1
Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks. However, LLMs are also prone to generate nonfactual content. Uncertainty Quantificat…
Handling Ontology Gaps in Semantic Parsing
Andrea Bacciu, Marco Damonte, Marco Basaldella +1
The majority of Neural Semantic Parsing (NSP) models are developed with the assumption that there are no concepts outside the ones such models can represent with their target symbo…