◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Roman Belaire

3 papers hereh-index 27 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2026

Robust Critics: Defending LLMs Against Multi-Turn Attacks

Roman Belaire, Arunesh Sinha, Pradeep Varakantham

When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM saf…

cs.LG2025

Automatic LLM Red Teaming

Roman Belaire, Arunesh Sinha, Pradeep Varakantham

Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prom…

cs.LG2025

On Minimizing Adversarial Counterfactual Error in Adversarial RL

Roman Belaire, Arunesh Sinha, Pradeep Varakantham

Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge in…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.