collaborators

5 papers

cs.RO2026

REBAR: Reference Ethical Benchmark for Autonomy Readiness

Jonathan Diller, David Barnes, Rebekah Bogdanoff +14

As autonomous systems grow more advanced, objective metrics to evaluate their ethical and legal compliance are critical for informing end users of their limitations and ensuring ac…

cs.AI2025

Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping

Dena Mujtaba, Brian Hu, Anthony Hoogs +1

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments.…

cs.CR2025

Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection

Tharindu Kumarage, Cameron Johnson, Jadie Adams +7

The rapid advancement of conversational agents, particularly chatbots powered by Large Language Models (LLMs), poses a significant risk of social engineering (SE) attacks on social…

cs.CL2025

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

Jadie Adams, Brian Hu, Emily Veenhuis +5

Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can on…

cs.CL2025

ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making

Bharadwaj Ravichandran, David Joy, Paul Elliott +6

Large language models (LLMs) are increasingly being used as decision aids. However, users have diverse values and preferences that can affect their decision-making, which requires…