2 papers
cs.AI2026
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
Guoyao Yu, Xiaoqing Sun, Ziqi Huang +13
Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order…
cs.CR2026
Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios
Lixun Ma, Ruolong Ma, Bei Wang +4
Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often re…