activity
20242026
most citedAutomatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks

1 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SE2026★ 1 cited

Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls

Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4

Large Language Models are increasingly deployed as judges (LaaJ) in code generation pipelines. While attractive for scalability, LaaJs tend to overlook domain specific issues raisi…

cs.SE2025

Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes

Ora Nova Fandina, Gal Amram, Eitan Farchi +6

Application modernization in legacy languages such as COBOL, PL/I, and REXX faces an acute shortage of resources, both in expert availability and in high-quality human evaluation d…

cs.SE2025

Automated Validation of LLM-based Evaluators for Software Engineering Artifacts

Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4

Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evalua…

cs.CL2025

LaajMeter: A Framework for LaaJ Evaluation

Samuel Ackerman, Gal Amram, Ora Nova Fandina +5

Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, c…

cs.SE2025

Quality Evaluation of COBOL to Java Code Transformation

Shmulik Froimovich, Raviv Gal, Wesam Ibraheem +1

We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in eva…

cs.SE2024★ 1 cited

Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks

Eitan Farchi, Shmulik Froimovich, Rami Katan +1

LLMs can be used in a variety of code related tasks such as translating from one programming language to another, implementing natural language requirements and code summarization.…