2 papers
cs.LG2026
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Gaojie Jin, Yong Tao, Lijia Yu +1
Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relying on the assumption that the…
cs.CV2025
REOBench: Benchmarking Robustness of Earth Observation Foundation Models
Xiang Li, Yong Tao, Siyuan Zhang +7
Earth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplo…