3 papers
cs.CL2026
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Haolin Chen, Deon Metelski, Leon Qi +30
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large l…
cs.LG2026
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
Youkang Wang, Jian Wang, Rubing Chen +3
Sequential scaling is a prominent inference-time scaling paradigm, yet its performance improvements are typically modest and not well understood, largely due to the prevalence of h…
cs.LG2025
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
Youkang Wang, Jian Wang, Rubing Chen +3
Test-time policy optimization enables large language models (LLMs) to adapt to distribution shifts by leveraging feedback from self-generated rollouts. However, existing methods re…