2 papers
cs.LG2026
Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
Dingyao Yu, Tong Zhang, Yutao Mou +3
LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate infe…
cs.CR2026
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Yutao Mou, Pengfei Yang, Zhe Yin +6
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely re…