activity
20192025
most citedSearch for and with inclusive tagging

44 citations · 589 across the 92 of their papers we have counts for

collaborators

156 papers

cs.SE2025

AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?

Ori Press, Brandon Amos, Haoyu Zhao +21

Despite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming…

cs.SE2025

SWE-smith: Scaling Data for Software Engineering Agents

John Yang, Kilian Lieret, Carlos E. Jimenez +7

Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point. Existing datasets are small, with at most 1,00…

cs.CL2024★ 2 cited

SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

John Yang, Carlos E. Jimenez, Alex L. Zhang +10

Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which…

cs.AI2024

EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities

Talor Abramovich, Meet Udeshi, Minghao Shao +13

Although language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited.…

cs.AI2024★ 3 cited

SciCode: A Research Coding Benchmark Curated by Scientists

Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang +27

Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evalua…

hep-ex2024★ 2 cited

Study of at Belle

Belle Collaboration, Z. S. Stottler, T. K. Pedlar +179

We report a study of the hadronic transitions , with , using mesons recorded by the Belle detector. We present the…