2 papers
cs.AI2025
Unbiased Evaluation of Large Language Models from a Causal Perspective
Meilin Chen, Jian Tian, Liang Ma +3
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation o…
cs.CV2025
The Fourth Monocular Depth Estimation Challenge
Anton Obukhov, Matteo Poggi, Fabio Tosi +54
This paper presents the results of the fourth edition of the Monocular Depth Estimation Challenge (MDEC), which focuses on zero-shot generalization to the SYNS-Patches benchmark, a…