7 papers
Evaluating Language Model Agency through Negotiations
Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4
We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
Dimitrios Rontogiannis, Maxime Peyrard, Nicolas Baldwin +3
Standard single-turn, static benchmarks fall short in evaluating the nuanced capabilities of Large Language Models (LLMs) on complex tasks such as software engineering. In this wor…
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
Clément Dumas, Chris Wendler, Veniamin Veselovsky +2
A central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In th…
Agentic AI: The Era of Semantic Decoding
Maxime Peyrard, Martin Josifoski, Robert West
Recent work demonstrated great promise in the idea of orchestrating collaborations between LLMs, human input, and various tools to address the inherent limitations of LLMs. We prop…
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
Veniamin Veselovsky, Berke Argin, Benedikt Stroebl +5
Just as humans display language patterns influenced by their native tongue when speaking new languages, LLMs often default to English-centric responses even when generating in othe…
Controlling Latent Diffusion Using Latent CLIP
Jason Becker, Chris Wendler, Peter Baylies +2
Instead of performing text-conditioned denoising in the image domain, latent diffusion models (LDMs) operate in latent space of a variational autoencoder (VAE), enabling more effic…