6 papers
Structured Thoughts For Improved Reasoning And Context Pruning
Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu +5
Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured T…
Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
Chia-Hsuan Lee, Zelei Cheng, Yu Wang +4
On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradie…
Critique-Guided Distillation for Robust Reasoning via Refinement
Berkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar +2
Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization. While c…
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
Chia-Hsuan Lee, Mingyang Zhou, Renkun Ni +6
Preference optimization methods such as DPO and KTO are widely used for aligning language models, yet little is understood about what properties of preference data drive downstream…
DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
Sahana Ramnath, Nima Chitsazan, Mingyang Zhou +8
Dialogues are a predominant mode of communication for humans, and it is immensely helpful to have automatically generated summaries of them (e.g., to revise key points discussed in…