2 citations · 2 across the 7 of their papers we have counts for
5 papers · 1 filter
CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions
Sherin Muckatira, Jesse Geneson, Slava Gerovitch +3
Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step…
PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges
Swastik Roy, Rajkumar Pujari, Tharindu Kumarage +5
LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be `…
Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition
Prasoon Goyal, Sattvik Sahai, Michael Johnston +14
Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Com…
Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
Tharindu Kumarage, Lisa Bauer, Yao Ma +7
As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks w…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…