5 papers
VERA: Variational Inference Framework for Jailbreaking Large Language Models
Anamika Lochab, Lu Yan, Patrick Pynadath +2
The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabilities in real-world settings. Without…
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
Qilin Liao, Anamika Lochab, Ruqi Zhang
Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplored vulnerabilities. Existing multimoda…
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Bolian Li, Yifan Wang, Yi Ding +3
Reinforcement learning (RL) has enabled complex reasoning abilities in large language models (LLMs). However, most RL algorithms suffer from performance saturation, preventing cont…
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
Anamika Lochab, Bolian Li, Ruqi Zhang
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often suffers from reduced multi-sa…
Energy-Based Reward Models for Robust Language Model Alignment
Anamika Lochab, Ruqi Zhang
Reward models (RMs) are essential for aligning Large Language Models (LLMs) with human preferences. However, they often struggle with capturing complex human preferences and genera…