activity
20222026
most citedIs ChatGPT a Good Causal Reasoner? A Comprehensive Evaluation

5 citations · 5 across the 11 of their papers we have counts for

collaborators

11 papers

cs.AI2026

Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration

Yang Zhao, Yangou Ouyang, Xiao Ding +8

While Hybrid Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become the standard paradigm for training LLM agents, effective mechanisms for data allocation…

cs.LG2026

MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization

Yang Zhao, Hepeng Wang, Xiao Ding +8

Group-Relative Policy Optimization (GRPO) has emerged as an efficient paradigm for aligning Large Language Models (LLMs), yet its efficacy is primarily confined to domains with ver…

cs.CL2025

Com: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models

Kai Xiong, Xiao Ding, Yixin Cao +7

Large language models (LLMs) have mastered abundant simple and explicit commonsense knowledge through pre-training, enabling them to achieve human-like performance in simple common…

cs.CL2025

CrossICL: Cross-Task In-Context Learning via Unsupervised Demonstration Transfer

Jinglong Gao, Xiao Ding, Lingxiao Zou +2

In-Context Learning (ICL) enhances the performance of large language models (LLMs) with demonstrations. However, obtaining these demonstrations primarily relies on manual effort. I…

cs.CL2025

ExpeTrans: LLMs Are Experiential Transfer Learners

Jinglong Gao, Xiao Ding, Lingxiao Zou +3

Recent studies provide large language models (LLMs) with textual task-solving experiences via prompts to improve their performance. However, previous methods rely on substantial hu…

cs.CL2025

Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection

Yang Zhao, Li Du, Xiao Ding +10

Large language models (LLMs) have shown great potential across various industries due to their remarkable ability to generalize through instruction tuning. However, the limited ava…