2 papers
cs.LG2026
SLIME: Stabilized Likelihood Implicit Margin Enforcement for Preference Optimization
Maksim Afanasyev, Illarion Iov
Direct preference optimization methods have emerged as a computationally efficient alternative to Reinforcement Learning from Human Feedback (RLHF) for aligning Large Language Mode…
cs.LG2025
Selective Adversarial Attacks on LLM Benchmarks
Ivan Dubrovsky, Anastasia Orlova, Illarion Iov +3
Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Pr…