3 papers
cs.CL2026
AI-Assisted Scientific Assessment: A Case Study on Climate Change
Christian Buck, Levke Caesar, Michelle Chen Huebscher +13
The emerging paradigm of AI co-scientists focuses on tasks characterized by repeatable verification, where agents explore search spaces in 'guess and check' loops. This paradigm do…
cs.AI2025
CLINB: A Climate Intelligence Benchmark for Foundational Models
Michelle Chen Huebscher, Katharine Mach, Aleksandar StaniÄ +10
Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CL…
cs.CL2024
Assessing Large Language Models on Climate Information
Jannis Bulian, Mike S. Schäfer, Afra Amini +8
As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grou…