2 papers
cs.LG2026
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
Hainiu Xu, Italo Luis da Silva, Jiangnan Ye +7
Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective. While existing safety rese…
stat.ML2026
Mean Testing under Truncation beyond Gaussian
Yuhao Wang, Roberto Imbuzeiro Oliveira, Themis Gouleakis
We characterize the fundamental limits of high-dimensional mean testing under arbitrary truncation, where samples are drawn from the conditional distribution for…