7 citations · 9 across the 11 of their papers we have counts for
Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
Lu Yan, Xuan Chen, Xiangyu Zhang
Current coding-agent benchmarks usually pro- vide the full task specification upfront. Real research coding often does not: the intended system is progressively disclosed through i…
cs.SE2026
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
Xuan Chen, Lu Yan, Ruqi Zhang +1
Large Language Model (LLM) agents increasingly act through external tools, making their safety contingent on tool-call workflows rather than text generation alone. While recent ben…