3 papers
cs.SE2025
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evalua…
cs.CL2025
LaajMeter: A Framework for LaaJ Evaluation
Samuel Ackerman, Gal Amram, Ora Nova Fandina +5
Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, c…
cs.SE2025
Quality Evaluation of COBOL to Java Code Transformation
Shmulik Froimovich, Raviv Gal, Wesam Ibraheem +1
We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in eva…