4 papers
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
Yaowei Zheng, Richong Zhang, Shenxi Wu +5
We study finite-horizon continuous-time policy evaluation from discrete closed-loop trajectories under time-inhomogeneous dynamics. The target value surface solves a backward parab…
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
Josue Torres-Fonseca, Naihao Deng, Yinpei Dai +5
Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insuffic…
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks
Shenxi Wu, Haosong Zhang, Xingjian Ma +4
Deeper modern architectures are costly to train, making hyperparameter transfer preferable to expensive repeated tuning. Maximal Update Parametrization (P) helps explain why ma…
Arithmetic-Mean P for Modern Architectures: A Unified Learning-Rate Scale for CNNs and ResNets
Haosong Zhang, Shenxi Wu, Yichi Zhang +2
Choosing an appropriate learning rate remains a key challenge in scaling depth of modern deep networks. The classical maximal update parameterization (P) enforces a fixed per-l…