closed-loop generation 1evaluation vs optimization 1llm-as-a-judge 1structured document understanding 1table recognition 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
Donghwan Kim
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track di…
cs.CL2026
LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition
Donghwan Kim
The paper investigates using large language models as judges to guide iterative table recognition, finding that their evaluation scores are weak and do not reliably improve generat…