2 papers
cs.CL2026
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
Esma Balkır, Alice Pernthaller, Marco Basaldella +2
Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tas…
cs.CL2025
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
Sanhanat Sivapiromrat, Caiqi Zhang, Marco Basaldella +1
Recent studies have shown that Large Language Models (LLMs) are vulnerable to data poisoning attacks, where malicious training examples embed hidden behaviours triggered by specifi…