collaborators

26 papers

math.PR2026

Graph Causal Optimal Transport and Wasserstein Distances

Jan Obłój, Vlad Tuchilus

We study the graph causal optimal transport problem, a generalisation of the classical optimal transport problem in which the allowed couplings satisfy causal restrictions prescrib…

cs.LG2026

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

Zelei Cheng, Amritansh Mishra, Sambit Sahu +1

Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural f…

cs.CL2026

Structured Thoughts For Improved Reasoning And Context Pruning

Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu +5

Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured T…

cs.CL2026

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5

Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…

cs.CL2026

SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

Chia-Hsuan Lee, Zelei Cheng, Yu Wang +4

On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradie…

cs.AI2026

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Sangwoo Cho, Kushal Chawla, Pengshan Cai +4

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, an…