Showing cs.SEShow all
3 papers · 1 filter
cs.SE2026
What Survives the Next Model? Benchmarking LLM-Based Techniques Against Single-Prompts
Nahian Salsabil, Joy Saha, Simantika Bhattacharjee Dristi +4
The software engineering research community has enthusiastically embraced the integration of Large Language Models (LLMs) into complex techniques to solve a wide variety of tasks.…
cs.SE2026
STADA: Specification-based Testing for Autonomous Driving Agents
Joy Saha, Trey Woodlief, Sebastian Elbaum +1
Simulation-based testing has become a standard approach to validating autonomous driving agents prior to real-world deployment. A high-quality validation campaign will exercise an…
cs.SE2024
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
Mohammed Latif Siddiq, Simantika Dristi, Joy Saha +1
Large Language Models (LLMs) are gaining popularity among software engineers. A crucial aspect of developing effective code generation LLMs is to evaluate these models using a robu…