5 papers
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
Terry Yue Zhuo, Xiaolong Jin, Hange Liu +37
Crowdsourced model evaluation platforms, such as Chatbot Arena, enable real-time evaluation from human perspectives to assess the quality of model responses. In the coding domain,…
Identifying and Mitigating API Misuse in Large Language Models
Terry Yue Zhuo, Junda He, Jiamou Sun +4
API misuse in code generated by large language models (LLMs) presents a serious and growing challenge in software development, as although LLMs demonstrate impressive code generati…
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
Yanli Wang, Yanlin Wang, Suiquan Wang +8
Repository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source reposito…
An Empirical Study on Low-Code Programming using Traditional vs Large Language Model Support
Yongkun Liu, Jiachi Chen, Tingting Bi +6
Low-code programming (LCP) refers to programming using models at higher levels of abstraction, resulting in less manual and more efficient programming, and reduced learning effort…
FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
Jiachi Chen, Yiming Shen, Jiashuo Zhang +7
High-quality smart contract vulnerability datasets are critical for evaluating security tools and advancing smart contract security research. Two major limitations of current manua…