Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents
Phat T. Tran-Truong, Xuan-Bach Le
Large language model (LLM) agents increasingly operate as sequential software systems, but their reliability is often summarized by scalar benchmark metrics. Metrics such as pass$@…
cs.SE2024
Adversarial Attacks on Code Models with Discriminative Graph Patterns
Thanh-Dat Nguyen, Yang Zhou, Xuan Bach D. Le +2
Pre-trained language models of code are now widely used in various software engineering tasks such as code generation, code completion, vulnerability detection, etc. This, in turn,…