2 citations · 2 across the 3 of their papers we have counts for
4 papers
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer +2
The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-sca…
Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Prahaladh Chandrahasan, Jiahe Jin, Zhihan Zhang +9
Effectively evaluating deep research agents that autonomously search the web, analyze information, and generate reports remains a major challenge, particularly when it comes to ass…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
NeoQA: Evidence-based Question Answering with Generated News Events
Max Glockner, Xiang Jiang, Leonardo F. R. Ribeiro +2
Evaluating Retrieval-Augmented Generation (RAG) in large language models (LLMs) is challenging because benchmarks can quickly become stale. Questions initially requiring retrieval…