4 papers
COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova +3
Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one…
Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding
Shuyin Ouyang, Jie M. Zhang, Jingzhi Gong +7
Software architecture diagrams are important design artifacts for communicating system structure, behavior, and data organization throughout the software development lifecycle. Alt…
Software is infrastructure: failures, successes, costs, and the case for formal verification
Giovanni Bernardi, Adrian Francalanza, Marco Peressotti +1
In this chapter we outline the role that software has in modern society, along with the staggering costs of poor software quality. To lay this bare, we recall the costs of some of…
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
Ying Xiao, Jie Huang, Ruijuan He +6
Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, unde…