4 papers
Stability-Aware Prompt Optimization for Clinical Data Abstraction
Arinbjörn Kolbeinsson, Daniel Timbie, Sajjan Narsinghani +1
Large language models used for clinical abstraction are sensitive to prompt wording, yet most work treats prompts as fixed and studies uncertainty in isolation. We argue these shou…
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Composable Interventions for Language Models
Arinbjorn Kolbeinsson, Kyle O'Brien, Tianjin Huang +8
Test-time interventions for language models can enhance factual accuracy, mitigate harmful outputs, and improve model efficiency without costly retraining. But despite a flood of n…
Adversarial Negotiation Dynamics in Generative Language Models
Arinbjörn Kolbeinsson, Benedikt Kolbeinsson
Generative language models are increasingly used for contract drafting and enhancement, creating a scenario where competing parties deploy different language models against each ot…