24 citations · 24 across the 2 of their papers we have counts for
3 papers
cs.CL2025
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI, :, Sandhini Agarwal +124
We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…
cs.CL2025★ 24 cited
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks +9
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations…
cs.SE2024
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
Foivos Tsimpourlas, Chao Peng, Carlos Rosuero +2
The Go programming language has gained significant traction for developing software, especially in various infrastructure systems. Nonetheless, concurrency bugs have become a preva…