activity
20242026
collaborators

8 papers

cs.CL2026

Post-Training Language Models for Crosslingual Consistency

Tianyu Liu, Jirui Qi, Mrinmaya Sachan +3

Language models often respond inconsistently to translation-equivalent prompts across languages, undermining the reliability of multilingual systems. To quantify this, we give an i…

cs.IR2026

How role-play shapes relevance judgment in zero-shot LLM rankers

Yumeng Wang, Jirui Qi, Catherine Chen +2

Large Language Models (LLMs) have emerged as promising zero-shot rankers, but their performance is highly sensitive to prompt formulation. In particular, role-play prompts, where t…

cs.CL2025

On the Consistency of Multilingual Context Utilization in Retrieval-Augmented Generation

Jirui Qi, Raquel Fernández, Arianna Bisazza

Retrieval-augmented generation (RAG) with large language models (LLMs) has demonstrated strong performance in multilingual question-answering (QA) tasks by leveraging relevant pass…

cs.CL2025

When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy

Jirui Qi, Shan Chen, Zidi Xiong +3

Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studi…

cs.CL2025

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Jirui Qi, Raquel Fernández, Arianna Bisazza

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. W…

cs.CL2025

Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation

Tianyu Liu, Jirui Qi, Paul He +3

Recent work suggests that large language models enhanced with retrieval-augmented generation are easily influenced by the order, in which the retrieved documents are presented to t…