4 papers
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
Gabriel Poesia, Simon Henniger, Tzu-Han Hsu +2
The cost of producing code is rapidly diminishing with increasingly capable AI agents, while quality assurance of generated programs has not kept pace. Formal verification provides…
Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains
Roy Rinberg, Annabelle Michael Carrell, Simon Henniger +2
We study the compression of LLM-generated text across lossless and lossy regimes, characterizing a compression-compute frontier where more compression is possible at the cost of mo…
Vision-Language Models Suppress Female Representations Under Ambiguous Input
Arnau Marin-Llobet, Simon Henniger, Mahzarin R. Banaji
Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous i…
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
Simon Henniger, Gabriel Poesia
Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is highly expensive, especially in re…