3 papers
cs.IR2026
How role-play shapes relevance judgment in zero-shot LLM rankers
Yumeng Wang, Jirui Qi, Catherine Chen +2
Large Language Models (LLMs) have emerged as promising zero-shot rankers, but their performance is highly sensitive to prompt formulation. In particular, role-play prompts, where t…
cs.CV2025
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts
Kiyoon Kim, Shreyank N Gowda, Panagiotis Eustratiadis +2
Despite recent advances in video action recognition achieving strong performance on existing benchmarks, these models often lack robustness when faced with natural distribution shi…
cs.CL2024
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
Xinyi Chen, Baohao Liao, Jirui Qi +4
Following multiple instructions is a crucial ability for large language models (LLMs). Evaluating this ability comes with significant challenges: (i) limited coherence between mult…