4 papers
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Large Language Models are increasingly deployed as judges (LaaJ) in code generation pipelines. While attractive for scalability, LaaJs tend to overlook domain specific issues raisi…
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
Itay Dreyfuss, Antonio Abu Nassar, Samuel Ackerman +5
Large Language Model (LLM)-based code assistants have emerged as a powerful application of generative AI, demonstrating impressive capabilities in code generation and comprehension…
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
Ora Nova Fandina, Gal Amram, Eitan Farchi +6
Application modernization in legacy languages such as COBOL, PL/I, and REXX faces an acute shortage of resources, both in expert availability and in high-quality human evaluation d…
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evalua…