3 papers
cs.CL2026
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
Esma Balkır, Alice Pernthaller, Marco Basaldella +2
Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tas…
cs.CL2025
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
Sanhanat Sivapiromrat, Caiqi Zhang, Marco Basaldella +1
Recent studies have shown that Large Language Models (LLMs) are vulnerable to data poisoning attacks, where malicious training examples embed hidden behaviours triggered by specifi…
cs.CL2024
LUQ: Long-text Uncertainty Quantification for LLMs
Caiqi Zhang, Fangyu Liu, Marco Basaldella +1
Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks. However, LLMs are also prone to generate nonfactual content. Uncertainty Quantificat…