21 papers
Multi-Turn On-Policy Distillation with Prefix Replay
Baohao Liao, Hanze Dong, Christof Monz +3
The paper introduces ReOPD, a method that reuses pre‑collected teacher trajectories as replayed prefixes to train LLM agents without costly new environment interactions, improving…
Fractured Chain-of-Thought Reasoning
Baohao Liao, Hanze Dong, Yuhui Xu +4
Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference…
On the Limits of Model Merging for Multilinguality in Pre-Training
Seth Aycock, Fedor Vitiugin, Aleksandr Umnov +2
Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-specific model merging. In this w…
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
Natalia Bila, Kata Naszádi, Kata Naszádi +2
We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions u…
Self-Hinting Language Models Enhance Reinforcement Learning
Baohao Liao, Hanze Dong, Xinxing Xu +2
Group Relative Policy Optimization (GRPO) has recently emerged as a practical recipe for aligning large language models with verifiable objectives. However, under sparse terminal r…
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
Xinlan Yan, Di Wu, Yibin Lei +2
In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset designed for benchmarking large language models (LLMs) in fine-grained clinical specialties.…