15 papers
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations
Xinshun Feng, Ziqi Miao, Lijun Li +1
Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical and the underlying knowledge is densely interconnected. In such se…
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Pengyu Zhu, Lijun Li, Longju Yang +2
Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist a…
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Chunxiao Li, Yuan Xiong, Lijun Li +4
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely…
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
Zhenhua Liu, Lijun Li, Ruizhe Chen +5
While guided decoding, especially value-guided methods, has emerged as a cost-effective alternative for controlling language model outputs without re-training models, its effective…
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model…
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
Yuexiao Liu, Lijun Li, Xingjun Wang +1
Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have gained significant attention due to their objective and verifiable reward signals, demonstrating s…