◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Benjamin Plaut

8 papers hereh-index 9523 citations24 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author3
  • last author4

Across the 8 of 8 papers where every author was matched, so the position is known.

fields
  • cs.LG6
  • cs.CL2

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing 2026Show all

4 papers · 1 filter

cs.CL2026

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen +1

As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is w…

cs.LG2026

Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards

Sarah Liaw, Benjamin Plaut

In high-stakes AI applications, even a single action can cause irreparable damage. However, nearly all of sequential decision-making theory assumes that all errors are recoverable…

cs.LG2026

Safety Training May Persist Through Helpfulness Optimization in LLM Agents

Benjamin Plaut

Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step…

cs.LG2026

YRC-Bench: A Benchmark for Learning to Coordinate with Experts

Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh +1

When deployed in the real world, AI agents will inevitably face challenges that exceed their individual capabilities. A critical component of AI safety is an agent's ability to rec…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.