Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Learning to Detect UI Principle Violations via Reinforcement Learning
Nishi Mehta, Swathi Alse, Himani Kumavat +3
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional correctness. A generated interf…
cs.CL2026
Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers
Guoxin Yu, Chulun Zhou, Lemao Liu +7
Evaluating the quality of model responses remains challenging in generative tasks with long-form answers, as the expected answers usually contain multiple semantically distinct yet…