2 papers
stat.ME2026
Heterogeneous Judge-Aware Ranking with Sensitivity, Disagreement, and Confidence
Shibo Yu, Yingzhou Wang, Yan Chen +2
Pairwise comparisons from multiple judges are central to large language model evaluation and preference modeling, yet standard ranking pipelines often pool judgments into a single…
cs.SE2025
DesignCoder: Hierarchy-Aware and Self-Correcting UI Code Generation with Large Language Models
Yunnong Chen, Shixian Ding, YingYing Zhang +4
Multimodal large language models (MLLMs) have streamlined front-end interface development by automating code generation. However, these models also introduce challenges in ensuring…