2 papers
cs.LG2026
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
Sheshansh Agrawal, Thien Hang Nguyen, Douwe Kiela
Selecting the top from items via expensive -wise comparisons is central to settings ranging from LLM-based document reranking to crowdsourced evaluation and tournament d…
cs.LG2026
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
Nick Ferguson, Josh Pennington, Narek Beghian +4
Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed…