10 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.CL2025★ 10 cited
AI-based Clinical Decision Support for Primary Care: A Real-World Study
Robert Korom, Sarah Kiptinness, Najib Adan +15
We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya,…
cs.AI2024★ 2 cited
Rule Based Rewards for Language Model Safety
Tong Mu, Alec Helyar, Johannes Heidecke +7
Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…
cs.CR2024★ 8 cited
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace, Kai Xiao, Reimar Leike +3
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…