12 papers · 1 filter
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning
Zheyuan Zhang, Manqing Mao, Hong Wang +8
Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of ro…
SupraBench: A Benchmark for Supramolecular Chemistry
Tianyi Ma, Yijun Ma, Zehong Wang +6
Supramolecular chemistry, which includes the study of non-covalent host-guest assemblies, has advanced various applications. However, designing host-guest systems remains time-cons…
ProPlay: Procedural World Models for Self-Evolving LLM Agents
Yijun Ma, Zehong Wang, Yiyang Li +5
Self-evolving agents are expected to improve through interaction without external supervision, but this remains difficult in partially observable environments where agents must exp…
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
Zheyuan Zhang, Kaiwen Shi, Han Bao +3
Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but l…
Hypergraph Pattern Machine: Compositional Tokenization for Higher-Order Interactions
Kyrie Zhao, Zehong Wang, Tianyi Ma +5
Hypergraphs model higher-order relations that drive real-world decisions, from drug prescriptions to recommendations. A central structural signal in such data, beyond what pairwise…
A Survey of Weight Space Learning: Understanding, Representation, and Generation
Xiaolong Han, Zehong Wang, Bo Zhao +8
Neural network weights are typically viewed as the end product of training, while most deep learning research focuses on data, features, and architectures. However, recent advances…