5 papers · 1 filter
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +5
Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire inc…
Learning from Reasoning Failures via Synthetic Data Generation
Gabriela Ben Melech Stan, Estelle Aflalo, Avinash Madasu +2
Training models on synthetic data has emerged as an increasingly important strategy for improving the performance of generative AI. This approach is particularly helpful for large…
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu +2
Are generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learning a world model from which sequences are generated one token at a tim…
DPO Learning with LLMs-Judge Signal for Computer Use Agents
Man Luo, David Cobbley, Xin Su +4
Computer use agents (CUA) are systems that automatically interact with graphical user interfaces (GUIs) to complete tasks. CUA have made significant progress with the advent of lar…
FastRM: An efficient and automatic explainability framework for multimodal generative models
Gabriela Ben-Melech Stan, Estelle Aflalo, Man Luo +5
Large Vision Language Models (LVLMs) have demonstrated remarkable reasoning capabilities over textual and visual inputs. However, these models remain prone to generating misinforma…