activity
20192025
most citedMultimodal Knowledge Alignment with Reinforcement Learning

18 citations · 58 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

31 papers · 1 filter

cs.CL20252 cited

Olmo 3

Team Olmo, :, Allyson Ettinger +66

We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long-context reasoning, function…

cs.CL2025

SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…

cs.CL20248 cited

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison, Valentina Pyatkin +20

Language model post-training is applied to refine behaviors and unlock new skills across a wide range of recent language models, but open recipes for applying these techniques lag…

cs.CL20242 cited

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

Yuling Gu, Oyvind Tafjord, Hyunwoo Kim +4

Large language models (LLMs) are increasingly tested for a "Theory of Mind" (ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at expli…

cs.CL20245 cited

WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries

Wenting Zhao, Tanya Goyal, Yu Ying Chiu +8

While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that t…

cs.CL20244 cited

WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu +6

We introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024…