collaborators

40 papers

cs.CR2026

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Yutao Mou, Pengfei Yang, Zhe Yin +6

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely re…

cs.AI2026

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Yidong Wang, Yan Zhan, Ziteng Feng +16

Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-spe…

cs.LG2026

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

Dingyao Yu, Tong Zhang, Yutao Mou +3

LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate infe…

cs.IR2026

SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval

Yuxiao Luo, Da Li, Mingjie Zhang +3

LLM-based retrievers have become a fundamental component of modern information retrieval systems. The paradigm of "rewrite-then-retriev" introduces explicit reasoning before retrie…

cs.CL2026

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics

Mengyuan Sun, Yu Li, Zhuohao Yu +2

Rubric-based evaluation is a promising paradigm for judging large language model (LLM) outputs, yet self-generated rubrics lag human-annotated criteria on hard instances. We argue…

cs.SE2026

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation

Zhengran Zeng, Ruikai Shi, Keke Han +7

Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Languag…