6 papers
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
Terry Yue Zhuo, Xiaolong Jin, Hange Liu +37
Crowdsourced model evaluation platforms, such as Chatbot Arena, enable real-time evaluation from human perspectives to assess the quality of model responses. In the coding domain,…
LLMAID: Identifying AI Capabilities in Android Apps with LLMs
Pei Liu, Terry Zhuo, Jiawei Deng +7
Recent advancements in artificial intelligence (AI) and its widespread integration into mobile software applications have received significant attention, highlighting the growing p…
PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
Pei Liu, Terry Zhuo, Jiawei Deng +4
The rapid emergence of pretrained models (PTMs) has attracted significant attention from both Deep Learning (DL) researchers and downstream application developers. However, selecti…
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
Junda He, Jieke Shi, Terry Yue Zhuo +5
The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks like code generation, producing a massive volume of software artifacts…
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
Xiaoxue Ren, Penghao Jiang, Kaixin Li +6
Web applications are prime targets for cyberattacks as gateways to critical services and sensitive data. Traditional penetration testing is costly and expertise-intensive, making i…
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
Haowei Quan, Junjie Wang, Xinzhe Li +3
In the rapidly evolving software development landscape, Python stands out for its simplicity, versatility, and extensive ecosystem. Python packages, as units of organization, reusa…