Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
Yang Zhao, Yangou Ouyang, Xiao Ding +8
While Hybrid Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become the standard paradigm for training LLM agents, effective mechanisms for data allocation…
cs.AI2022
e-CARE: a New Dataset for Exploring Explainable Causal Reasoning
Li Du, Xiao Ding, Kai Xiong +2
Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can…