2 citations · 2 across the 2 of their papers we have counts for
4 papers
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
Adarsh Kumarappan, Pareesa Ameneh Golnari, Wen Wen +5
DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,800 evaluation instances across six pro…
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot
Fangchen Song, Ashish Agarwal, Wen Wen
Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for developer productivity and participation i…
Role Play: Learning Adaptive Role-Specific Strategies in Multi-Agent Interactions
Weifan Long, Wen Wen, Peng Zhai +1
Zero-shot coordination problem in multi-agent reinforcement learning (MARL), which requires agents to adapt to unseen agents, has attracted increasing attention. Traditional approa…
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
Emman Haider, Daniel Perez-Becker, Thomas Portet +28
Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…