3 papers
cs.CY2026
AI Security Priorities: A Field-Wide Agenda
Gil Gekker, Rachel Steratore, Everett Smith +9
As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen…
cs.LG2025
Leveraging LLM Inconsistency to Boost Pass@k Performance
Uri Dalal, Meirav Segal, Zvika Ben-Haim +2
Large language models (LLMs) achieve impressive abilities in numerous domains, but exhibit inconsistent performance in response to minor input changes. Rather than view this as a d…
cs.LG2025
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
Gil Gekker, Meirav Segal, Dan Lahav +1
Following the rapid increase in Artificial Intelligence (AI) capabilities in recent years, the AI community has voiced concerns regarding possible safety risks. To support decision…