3 papers
cs.AI2026
Probing RLVR training instability through the lens of objective-level hacking
Yiming Dong, Kun Fu, Haoyu Li +5
Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the trai…
cs.LG2026
MetaKube: An Experience-Aware LLM Framework for Kubernetes Failure Diagnosis
Wei Sun, Ting Wang, Xinran Tian +4
Existing LLM-based Kubernetes diagnostic systems cannot learn from operational experience, operating on static knowledge bases without improving from past resolutions. We present M…
cs.CL2025
From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
Haoyu Li, Xuhong Li, Yiming Dong +1
Dataset diversity plays a pivotal role for the successful training of many machine learning models, particularly in the supervised fine-tuning (SFT) stage of large language model (…