4 papers
Configurable Preference Tuning with Rubric-Guided Synthetic Data
Víctor Gallego
Models of human feedback for AI alignment, such as those underpinning Direct Preference Optimization (DPO), often bake in a singular, static set of preferences, limiting adaptabili…
Merging Improves Self-Critique Against Jailbreak Attacks
Victor Gallego
The robustness of large language models (LLMs) against adversarial manipulations, such as jailbreak attacks, remains a significant challenge. In this work, we propose an approach t…
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
Víctor Gallego
In this paper, we introduce \emph{refined Direct Preference Optimization} (rDPO), a method for improving the behavioral alignment of Large Language Models (LLMs) without the need f…
Fast Adaptation with Bradley-Terry Preference Models in Text-To-Image Classification and Generation
Victor Gallego
Recently, large multimodal models, such as CLIP and Stable Diffusion have experimented tremendous successes in both foundations and applications. However, as these models increase…