◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

V. Hebbar

5 papers hereh-index 3364 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3
  • last author2

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1
  • cs.CL1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.LGShow all

3 papers · 1 filter

cs.LG2026

Diffuse AI Control on Fuzzy Tasks

Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1

AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned wit…

cs.LG2026

Removing Sandbagging in LLMs by Training with Weak Supervision

Emil Ryd, Henning Bartsch, Julian Stastny +2

As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more cap…

cs.LG2025

Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Abhay Sheshadri, Aidan Ewart, Phillip Guo +8

Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.