6 papers
Governed Shared Memory for Multi-Agent LLM Systems
Yanki Margalit, Nurit Cohen-Inger, Erni Avram +2
Multi-agent LLM environments require robust mechanisms for shared knowledge management. This paper formalizes the fleet-memory problem and identifies four foundational failure mode…
PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
Yanki Margalit, Erni Avram, Ran Taig +2
Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly…
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
Nurit Cohen-Inger, Yehonatan Elisha, Bracha Shapira +2
Large language models (LLMs) often appear to excel on public benchmarks, but these high scores may mask an overreliance on dataset-specific surface cues rather than true language u…
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
Seffi Cohen, Niv Goldshlager, Nurit Cohen-Inger +2
Large Language Models (LLMs) have shown remarkable capabilities across various natural language processing tasks but often struggle to excel uniformly in diverse or complex domains…
FairTTTS: A Tree Test Time Simulation Method for Fairness-Aware Classification
Nurit Cohen-Inger, Lior Rokach, Bracha Shapira +1
Algorithmic decision-making has become deeply ingrained in many domains, yet biases in machine learning models can still produce discriminatory outcomes, often harming unprivileged…
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
Nurit Cohen-Inger, Seffi Cohen, Neomi Rabaev +2
As machine learning (ML) systems increasingly impact critical sectors such as hiring, financial risk assessments, and criminal justice, the imperative to ensure fairness has intens…