Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
IMPersona: Evaluating Individual Level LM Impersonation
Quan Shi, Carlos E. Jimenez, Stephen Dong +4
As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the charac…
cs.CL2024
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig +4
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-w…
cs.CL2024
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
John Yang, Carlos E. Jimenez, Alex L. Zhang +10
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which…