4 papers
The Effect of Multi-Lingual and Keyword Adversarial Injection on LLM Relevance Judgment
Nguyen Khoi Vo, Duy Duong Tuong, Oleg Zendel +1
Large language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation rem…
RMIT-ADM+S at the MMU-RAG NeurIPS 2025 Competition
Kun Ran, Marwah Alaofi, Danula Hettiachchi +9
This paper presents the award-winning RMIT-ADM+S system for the Text-to-Text track of the NeurIPS~2025 MMU-RAG Competition. We introduce Routing-to-RAG (R2RAG), a research-focused…
RMIT-ADM+S at the SIGIR 2025 LiveRAG Challenge
Kun Ran, Shuoqi Sun, Khoi Nguyen Dinh Anh +2
This paper presents the RMIT--ADM+S winning system in the SIGIR 2025 LiveRAG Challenge. Our Generation-Retrieval-Augmented Generation (G-RAG) approach generates a hypothetical answ…
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
Laura Dietz, Oleg Zendel, Peter Bailey +6
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empi…