26 papers
Graph Causal Optimal Transport and Wasserstein Distances
Jan Obłój, Vlad Tuchilus
We study the graph causal optimal transport problem, a generalisation of the classical optimal transport problem in which the allowed couplings satisfy causal restrictions prescrib…
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
Zelei Cheng, Amritansh Mishra, Sambit Sahu +1
Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural f…
Structured Thoughts For Improved Reasoning And Context Pruning
Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu +5
Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured T…
Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
Chia-Hsuan Lee, Zelei Cheng, Yu Wang +4
On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradie…
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Sangwoo Cho, Kushal Chawla, Pengshan Cai +4
Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, an…