activity
20242026
collaborators

10 papers

cs.SE2026

When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models

Shenyu Zheng, Ximing Dong, Xiaoshuang Liu +6

As Large Language Models (LLMs) achieve breakthroughs in complex reasoning, Codeforces-based Elo ratings have emerged as a prominent metric for evaluating competitive programming c…

cs.SE2025

SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints

Zhiyu Fan, Kirill Vasilevski, Dayi Lin +6

The advancement of large language models (LLMs) and code agents has demonstrated significant potential to assist software engineering (SWE) tasks, such as autonomous issue resoluti…

cs.SE2025

RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale

Zhilong Chen, Chengzong Zhao, Boyuan Chen +9

Training software engineering (SWE) LLMs is bottlenecked by expensive infrastructure, inefficient evaluation pipelines, scarce training data, and costly quality control. We present…

cs.SE2025

SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation

Gustavo A. Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia +7

High-quality labeled datasets are crucial for training and evaluating foundation models in software engineering, but creating them is often prohibitively expensive and labor-intens…

cs.SE2025

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)

Kirill Vasilevski, Benjamin Rombaut, Gopi Krishnan Rajbahadur +10

Foundation Models (FMs) such as Large Language Models (LLMs) are reshaping the software industry by enabling FMware, systems that integrate these FMs as core components. In this KD…

cs.SE2025

Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement

Keheliya Gallaba, Ali Arabat, Dayi Lin +2

Foundation Models (FMs) have shown remarkable capabilities in various natural language tasks. However, their ability to accurately capture stakeholder requirements remains a signif…