3 papers
cs.CL2025
LaajMeter: A Framework for LaaJ Evaluation
Samuel Ackerman, Gal Amram, Ora Nova Fandina +5
Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, c…
cs.SE2025
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evalua…
cs.SE2025
Quality Evaluation of COBOL to Java Code Transformation
Shmulik Froimovich, Raviv Gal, Wesam Ibraheem +1
We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in eva…