8 papers
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Guangxiang Zhao, Qilong Shi, Xusen Xiao +13
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from…
ArborMem: Navigating Interaction States with Memory Forests
Zongwei Lv, Yuemeng Xu, Yilun Yao +8
Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions.…
When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment
Weihong Lin, Lin Sun, Xiangzheng Zhang
Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under…
SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills
Nizhang Li, Zonghao Ying, Xiangfan Wu +7
External skills extend the capabilities of large language model agents, but also introduce an execution-time attack surface: a skill that appears benign under inspection may reveal…
SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
Haowen Dai, Zonghao Ying, Wenfeng Li +10
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective c…
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
Dongdong Yang, Deyue Zhang, Zhao Liu +5
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly…