1 citations · 2 across the 6 of their papers we have counts for
6 papers
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Large Language Models are increasingly deployed as judges (LaaJ) in code generation pipelines. While attractive for scalability, LaaJs tend to overlook domain specific issues raisi…
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
Ora Nova Fandina, Gal Amram, Eitan Farchi +6
Application modernization in legacy languages such as COBOL, PL/I, and REXX faces an acute shortage of resources, both in expert availability and in high-quality human evaluation d…
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evalua…
LaajMeter: A Framework for LaaJ Evaluation
Samuel Ackerman, Gal Amram, Ora Nova Fandina +5
Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, c…
Quality Evaluation of COBOL to Java Code Transformation
Shmulik Froimovich, Raviv Gal, Wesam Ibraheem +1
We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in eva…
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
Eitan Farchi, Shmulik Froimovich, Rami Katan +1
LLMs can be used in a variety of code related tasks such as translating from one programming language to another, implementing natural language requirements and code summarization.…