2 papers
cs.AI2026
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Tianyou Wang, Chongyang Gao, Kezhen Chen +6
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the…
cs.CR2026
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
Chia-Yi Hsu, Chia-Mu Yu, Chun-Ying Huang +1
LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package installation commands. This c…