2 papers
cs.CL2026
Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
Himil Vasava, Ming Jiang
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they a…
cs.CV2026
Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?
Yun Li, Biao Yang, Peixi Wu +5
Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess representations through discrim…