13 papers
Mutate to Bypass: Autonomous Endpoint Evasion via Knowledge-Driven Multi-Agent Orchestration
Weifeng Yuan, Wenbo Guo, Qingyun Du +4
Public reports and open-source resources expose many EDR evasion techniques, but it remains unclear whether commercial Endpoint Detection and Response (EDR) systems can withstand t…
Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
Weifeng Yuan, Wenbo Guo, Feng Dong +2
The paper investigates how large language model agents often fabricate nonexistent skill names when asked to recommend and install skills, exposing a supply‑chain security risk, an…
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
Xian Wu, Kaijie Zhu, Ying Zhang +2
Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking. In LLM reasoning, existing works a…
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study
Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or in…
Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
Xiaojun Jia, Jie Liao, Simeng Qin +5
Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely o…
MalwarePT: A Binary-Level Foundation Model for Malware Analysis
Saastha Vasan, Yuzhou Nie, Kaie Chen +6
Automated malware analysis increasingly relies on machine learning, yet most existing methods remain task-specific and depend on handcrafted features or narrowly scoped models. Rec…