7 papers
A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
Jianlin Chen, Wenhui Chen, Ziyao Lin +1
LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct va…
Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B
Defu Lin, Wenhui Chen, Ziyao Lin +3
Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constraint through four resource walls: a percept…
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
Wenhui Chen, Jianlin Chen, Ziyao Lin +2
Language models are increasingly promoted from examinees to examiners: they write the test suites, answer keys, rubrics, and reward functions that define correctness for other syst…
Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable
Wenhui Chen, Jianlin Chen, Ziyao Lin +1
When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain…
The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
Wenhui Chen, Jianlin Chen, Ziyao Lin +1
The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality. We propose its sequel…
DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
Zhechao Wang, Yiming Zeng, Lufan Ma +4
Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current method…