activity
20212026
most citedDataPerf: Benchmarks for Data-Centric AI Development

51 citations · 265 across the 31 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Reflective Context Learning: Studying the Optimization Primitives of Context Space

Nikita Vassilyev, William Berrios, Ruowang Zhang +3

Generally capable agents must learn from experience in ways that generalize across tasks and environments. The fundamental problems of learning, including credit assignment, overfi…

cs.LG2026

ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction

Nick Ferguson, Josh Pennington, Narek Beghian +4

Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed…

cs.LG2026

BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs

Sheshansh Agrawal, Thien Hang Nguyen, Douwe Kiela

Selecting the top from items via expensive -wise comparisons is central to settings ranging from LLM-based document reranking to crowdsourced evaluation and tournament d…

cs.LG2025★ 2 cited

Great Models Think Alike and this Undermines AI Oversight

Shashwat Goel, Joschka Struber, Ilze Amanda Auzina +6

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these…

cs.LG2024★ 1 cited

Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment

Karel D'Oosterlinck, Winnie Xu, Chris Develder +5

Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes…

cs.LG2024★ 23 cited

KTO: Model Alignment as Prospect Theoretic Optimization

Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff +2

Kahneman & Tversky's tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-ave…