6 papers
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
Danlu Chen, Ka Sing He, Jiahe Tian +4
The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language…
Zephyrus: An Agentic Framework for Weather Science
Sumanth Varambally, Marshall Fisher, Jas Thakker +14
Foundation models for weather science are pre-trained on vast amounts of structured numerical data and outperform traditional weather forecasting systems. However, these models lac…
Studying the Soupability of Documents in State Space Models
Yasaman Jafari, Zixian Wang, Leon Bergen +1
We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document…
Single-Pass Document Scanning for Question Answering
Weili Cao, Jianyou Wang, Youze Zheng +5
Handling extremely large documents for question answering is challenging: chunk-based embedding methods often lose track of important global context, while full-context transformer…
Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation
Bohan Lyu, Yadi Cao, Duncan Watson-Parris +3
Large Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools c…
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
Veeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky +6
The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation fram…