4 papers · 1 filter
Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards
Bizhe Bai, Xinyue Wang, Peng Ye +1
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning, yet growing evidence indicates an exploration ceiling: it often reweights existing solution traces rat…
Transformer Is Inherently a Causal Learner
Xinyue Wang, Stephen Wang, Biwei Huang
We reveal that transformers trained in an autoregressive manner naturally encode time-delayed causal structures in their learned representations. When predicting future values in m…
ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering
Xinyue Wang, Yuheng Jia, Hui Liu +1
Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed under…
Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations
Yupei Yang, Biwei Huang, Fan Feng +3
General intelligence requires quick adaption across tasks. While existing reinforcement learning (RL) methods have made progress in generalization, they typically assume only distr…