2 papers
cs.AI2026
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
Yiyao Zhang, Diksha Goel, Hussain Ahmad +2
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accur…
cs.AI2026
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
Yiyao Zhang, Diksha Goel, Hussain Ahmad +2
A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement wit…