5 papers
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks
Yifan Chen, Haitao Li, Yiran Hu +6
As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…
Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation
Yongjie Wang, Xinyue Zhang, Kunhong Yao +4
Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such age…
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
Yige Xu, Yongjie Wang, Zizhuo Wu +3
Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear…
Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
Xurui Li, Kaisong Song, Rui Zhu +2
Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either i…
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
Yongjie Wang, Yue Yu, Kaisong Song +2
Large Language Models (LLMs) have enabled a wide range of applications through their powerful capabilities in language understanding and generation. However, as LLMs are trained on…