Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
Ora Nova Fandina, Leshem Choshen, Eitan Farchi +3
Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs,…
cs.AI2024
Generating Unseen Code Tests In Infinitum
Marcel Zalmanovici, Orna Raz, Eitan Farchi +1
Large Language Models (LLMs) are used for many tasks, including those related to coding. An important aspect of being able to utilize LLMs is the ability to assess their fitness fo…