11 papers
The Clinical Trial Pipeline Reveals the Next Wave of Artificial Intelligence in Healthcare: A Multidimensional Analysis of 8,532 Registered Studies
Lior Rokach
The prospective clinical evaluation of artificial intelligence in medicine has expanded rapidly, but the global AI clinical trial landscape remains incompletely characterized. We s…
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
Bronislav Sidik, Lior Rokach
Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over 72-hour operation windows d…
Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)
Assaf Gerner, Netta Madvil, Nadav Barak +11
Large Language Models (LLMs) augmented with Retrieval-Augmented Generation (RAG) techniques are revolutionizing applications across multiple domains, such as healthcare, finance, a…
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
Bronislav Sidik, Lior Rokach
Autonomous AI agents built on open-source runtimes such as OpenClaw expose every available tool to every session by default, regardless of the task. A summarization task receives t…
SHAPoint: Task-Agnostic, Efficient, and Interpretable Point-Based Risk Scoring via Shapley Values
Tomer D. Meirman, Bracha Shapira, Noa Dagan +1
Interpretable risk scores play a vital role in clinical decision support, yet traditional methods for deriving such scores often rely on manual preprocessing, task-specific modelin…
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
Nurit Cohen-Inger, Yehonatan Elisha, Bracha Shapira +2
Large language models (LLMs) often appear to excel on public benchmarks, but these high scores may mask an overreliance on dataset-specific surface cues rather than true language u…