3 papers
cs.SE2026
Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents
Liting Lin, Boxi Yu, Yuzhong Zhang +3
Conversational LLM agents can cause real-world harm when their internal workflows fail, such as completing a transaction without confirmation. Testing these state-dependent failure…
cs.CL2026
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
Boxi Yu, Yuzhong Zhang, Liting Lin +2
Large language models can still hallucinate in retrieval-augmented generation (RAG), producing claims that are unsupported by or conflict with the retrieved context. Detecting such…
cs.SE2026
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang +9
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals th…