3 citations · 3 across the 20 of their papers we have counts for
Showing 2026 · cs.CLShow all
3 papers · 2 filters
cs.CL2026
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
Chishui Chen, Yaoyou Fan, Te Sun +11
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn age…
cs.CL2026
Where to Place the Query? Unveiling and Mitigating Positional Bias in In-Context Learning for Diffusion LLMs via Decoding Dynamics
Zhengheng Li, Panrui Li, Xuyang Liu +1
While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remains largely unexplored. Unlike…
cs.CL2026
STDec: Spatio-Temporal Stability Guided Decoding for dLLMs
Yuzhe Chen, Jiale Cao, Xuyang Liu +3
Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a gl…