activity
20242026
collaborators

5 papers

cs.LG2026

MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation

Yutong Wang, Pengliang Ji, Chaoqun Yang +4

The LLM-as-a-Judge paradigm shows promise for evaluating generative content but lacks reliability in reasoning-intensive scenarios, such as programming. Inspired by recent advances…

cs.SE2026

Assessing the Impact of Requirement Ambiguity on LLM-based Function-Level Code Generation

Di Yang, Xinou Xie, Xiuwen Yang +7

Software requirement ambiguity is ubiquitous in real-world development, stemming from the inherent imprecision of natural language and the varying interpretations of stakeholders.…

cs.SE2025

Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision

Xu Lu, Weisong Sun, Yiran Zhang +4

Automated code generation has long been considered the holy grail of software engineering. The emergence of Large Language Models (LLMs) has catalyzed a revolutionary breakthrough…

cs.SE2025

Intention is All You Need: Refining Your Code from Your Intention

Qi Guo, Xiaofei Xie, Shangqing Liu +3

Code refinement aims to enhance existing code by addressing issues, refactoring, and optimizing to improve quality and meet specific requirements. As software projects scale in siz…

cs.DC2024

NebulaFL: Effective Asynchronous Federated Learning for JointCloud Computing

Fei Gao, Ming Hu, Zhiyu Xie +4

With advancements in AI infrastructure and Trusted Execution Environment (TEE) technology, Federated Learning as a Service (FLaaS) through JointCloud Computing (JCC) is promising t…