8 papers
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
Bardia Mohammadi, Lars Klein, Akhil Arora +1
Tool-augmented language agents speculatively issue likely future tool calls to hide latency, but those calls leak inferred user intent to external services before the agent commits…
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
Nearchos Potamitis, Vansh Ramani, Har Ashish Arora +3
Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeat…
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
Xavier Theimer-Lienhard, Mushtaha El-Amin, Fay Elhassan +5
Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Mos…
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
Bardia Mohammadi, Nearchos Potamitis, Lars Klein +2
LLM agents execute multi-step workflows that mutate external state through tools. Common orchestrators treat tool return as the settlement trigger, so faults, speculation, and conc…
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Fay Elhassan, David Sasu, Alexandra Kulinkina +2
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive O…
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
Yusuf Kesmen, Fay Elhassan, Jiayi Ma +7
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argu…