4 papers
PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation
Yangyang Feng, Zhuoyan Feng, Junlan Chen
On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollout also reveals how the student…
Two-way Evidence self-Alignment based Dual-Gated Reasoning Enhancement
Kexin Zhang, Junlan Chen, Daifeng Li +4
Large language models (LLMs) encounter difficulties in knowledge-intensive multi-step reasoning (KIMSR) tasks. One challenge is how to effectively extract and represent rationale e…
Can Competition Enhance the Proficiency of Agents Powered by Large Language Models in the Realm of News-driven Time Series Forecasting?
Yuxuan Zhang, Yangyang Feng, Daifeng Li +3
Multi-agents-based news-driven time series forecasting is considered as a potential paradigm shift in the era of large language models (LLMs). The challenge of this task lies in me…
Structuring Scientific Innovation: A Framework for Modeling and Discovering Impactful Knowledge Combinations
Junlan Chen, Kexin Zhang, Daifeng Li +3
The emergence of large language models offers new possibilities for structured exploration of scientific knowledge. Rather than viewing scientific discovery as isolated ideas or co…