Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
Input Relation Prompting for Metamorphic Testing on Query-Based Systems
Eng-Shen Tu, Shin-Jie Lee
Testing query-based systems (QBSs) presents significant challenges due to the absence of ground truth for validation and the extensive time and effort required for manual testing.…
cs.SE2026
OmniCode: A Benchmark for Evaluating Software Engineering Agents
Atharv Sonwane, Eng-Shen Tu, Wei-Chung Lu +11
LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require challenging benchmarks that can rigoro…