10 papers
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
Alexander Zadorojniy, Segev Wasserkrug, Eitan Farchi
Recently, using Large Language Models (LLMs) to generate optimization models from natural language descriptions has became increasingly popular. However, a major open question is h…
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich +4
Large Language Models are increasingly deployed as judges (LaaJ) in code generation pipelines. While attractive for scalability, LaaJs tend to overlook domain specific issues raisi…
Enhancing Formal Software Specification with Artificial Intelligence
Antonio Abu Nassar, Eitan Farchi
Formal software specification is known to enable early error detection and explicit invariants, yet it has seen limited industrial adoption due to its high notation overhead and th…
LaajMeter: A Framework for LaaJ Evaluation
Samuel Ackerman, Gal Amram, Ora Nova Fandina +5
Large Language Models (LLMs) are increasingly used as evaluators in natural language processing tasks, a paradigm known as LLM-as-a-Judge (LaaJ). The analysis of a LaaJ software, c…
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
Eitan Farchi, Kiran Nayak, Papia Ghosh Majumdar +1
Large Language Models (LLMs) are transforming Quality Engineering (QE) by automating the generation of artefacts such as requirements, test cases, and Behavior Driven Development (…
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
Ora Nova Fandina, Gal Amram, Eitan Farchi +6
Application modernization in legacy languages such as COBOL, PL/I, and REXX faces an acute shortage of resources, both in expert availability and in high-quality human evaluation d…