Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Reward Hacking as Equilibrium under Finite Evaluation
Jiacheng Wang, Jinbin Huang
We prove that under five minimal axioms -- multi-dimensional quality, finite evaluation, effective optimization, resource finiteness, and combinatorial interaction -- any optimized…
cs.AI2023
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
Jinbin Huang, Wenbin He, Liang Gou +2
Deep learning models are widely used in critical applications, highlighting the need for pre-deployment model understanding and improvement. Visual concept-based methods, while inc…