4 papers
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Large Language Models are increasingly deployed as judges (LaaJ) in code generation pipelines. While attractive for scalability, LaaJs tend to overlook domain specific issues raisi…
Evaluating perturbation robustness of generative systems that use COBOL code inputs
Samuel Ackerman, Wesam Ibraheem, Orna Raz +1
Systems incorporating large language models (LLMs) as a component are known to be sensitive (i.e., non-robust) to minor input variations that do not change the meaning of the input…
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
Ora Nova Fandina, Gal Amram, Eitan Farchi +6
Application modernization in legacy languages such as COBOL, PL/I, and REXX faces an acute shortage of resources, both in expert availability and in high-quality human evaluation d…
Quality Evaluation of COBOL to Java Code Transformation
Shmulik Froimovich, Raviv Gal, Wesam Ibraheem +1
We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in eva…