Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
Alexander Zadorojniy, Segev Wasserkrug, Eitan Farchi
Recently, using Large Language Models (LLMs) to generate optimization models from natural language descriptions has became increasingly popular. However, a major open question is h…
cs.AI2025
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
Ora Nova Fandina, Leshem Choshen, Eitan Farchi +3
Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs,…