◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Rajdeep Haldar

5 papers hereh-index 213 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • stat.ML1

identity via Semantic Scholar / OpenAlex

activity
20232026
collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

Rajdeep Haldar, Lantao Mei, Guang Lin +2

Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-preferred) distributions, yielding…

cs.LG2025

LLM Safety Alignment is Divergence Estimation in Disguise

Rajdeep Haldar, Ziyi Wang, Qifan Song +2

We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or…

cs.LG2024

Effect of Ambient-Intrinsic Dimension Gap on Adversarial Vulnerability

Rajdeep Haldar, Yue Xing, Qifan Song

The existence of adversarial attacks on machine learning models imperceptible to a human is still quite a mystery from a theoretical perspective. In this work, we introduce two not…

cs.LG2023

On Neural Network approximation of ideal adversarial attack and convergence of adversarial training

Rajdeep Haldar, Qifan Song

Adversarial attacks are usually expressed in terms of a gradient-based operation on the input data and model, this results in heavy computations every time an attack is generated.…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.