domain specialization 1large language model evaluation 1LoRA adapters 1model deferral 1reward benchmarking 1risk auditing 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation
Ye Chen, Weining Zhang
The paper studies whether LLM evaluators should be domain‑specialized in their weights or only in the decision rule that defers judgment, showing that sharing a common judge works…
cs.LG2026
GLM-5: from Vibe Coding to Agentic Engineering
GLM-5-Team, :, Aohan Zeng +184
We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (AR…