51 citations · 265 across the 31 of their papers we have counts for
8 papers · 1 filter
Reflective Context Learning: Studying the Optimization Primitives of Context Space
Nikita Vassilyev, William Berrios, Ruowang Zhang +3
Generally capable agents must learn from experience in ways that generalize across tasks and environments. The fundamental problems of learning, including credit assignment, overfi…
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
Nick Ferguson, Josh Pennington, Narek Beghian +4
Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed…
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
Sheshansh Agrawal, Thien Hang Nguyen, Douwe Kiela
Selecting the top from items via expensive -wise comparisons is central to settings ranging from LLM-based document reranking to crowdsourced evaluation and tournament d…
Great Models Think Alike and this Undermines AI Oversight
Shashwat Goel, Joschka Struber, Ilze Amanda Auzina +6
As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these…
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
Karel D'Oosterlinck, Winnie Xu, Chris Develder +5
Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes…
KTO: Model Alignment as Prospect Theoretic Optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff +2
Kahneman & Tversky's tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-ave…