26 citations · 37 across the 16 of their papers we have counts for
3 papers · 1 filter
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
Itay Dreyfuss, Antonio Abu Nassar, Samuel Ackerman +5
Large Language Model (LLM)-based code assistants have emerged as a powerful application of generative AI, demonstrating impressive capabilities in code generation and comprehension…
Evaluating perturbation robustness of generative systems that use COBOL code inputs
Samuel Ackerman, Wesam Ibraheem, Orna Raz +1
Systems incorporating large language models (LLMs) as a component are known to be sensitive (i.e., non-robust) to minor input variations that do not change the meaning of the input…
Statistical multi-metric evaluation and visualization of LLM system predictive performance
Samuel Ackerman, Eitan Farchi, Orna Raz +1
The evaluation of generative or discriminative large language model (LLM)-based systems is often a complex multi-dimensional problem. Typically, a set of system configuration alter…