3 papers
cs.AI2026
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Jing Chen, Yang Sun, Li Zhang +2
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmar…
cs.CL2026
Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement
Yuxiao Lu, Lin Xu, Yang Sun +2
Large language models (LLMs) aligned for safety often suffer from over-refusal, the tendency to reject seemingly toxic or benign prompts by misclassifying them as toxic. This behav…
cs.CR2025
SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
Haotian Xu, Qingsong Peng, Jie Shi +3
The rapid adoption of large language models (LLMs) in critical domains has spurred extensive research into their security issues. While input manipulation attacks (e.g., prompt inj…