Showing 2025Show all
2 papers · 1 filter
cs.SE2025
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
Itay Dreyfuss, Antonio Abu Nassar, Samuel Ackerman +5
Large Language Model (LLM)-based code assistants have emerged as a powerful application of generative AI, demonstrating impressive capabilities in code generation and comprehension…
stat.AP2025
Statistical multi-metric evaluation and visualization of LLM system predictive performance
Samuel Ackerman, Eitan Farchi, Orna Raz +1
The evaluation of generative or discriminative large language model (LLM)-based systems is often a complex multi-dimensional problem. Typically, a set of system configuration alter…