8 papers
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Juliusz Ziomek, William Bankes, Lorenz Wolf +3
We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, models must efficiently navigate Wi…
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Xiaohang Tang, Keyue Jiang, Che Liu +4
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likeli…
On-Policy Adversarial Flow Distillation for Autoregressive Video Generation
Yang Luo, Shengju Qian, Xiaohang Tang +4
Autoregressive video generators are attractive for streaming, long-horizon, and interactive applications, but distilling strong black-box teachers into causal students remains diff…
Robust Multi-Objective Controlled Decoding of Large Language Models
Seongho Son, William Bankes, Sangwoong Yoon +3
We introduce Robust Multi-Objective Decoding (RMOD), a novel inference-time algorithm that robustly aligns Large Language Models (LLMs) to multiple human objectives (e.g., instruct…
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
Xiaohang Tang, Rares Dolga, Sangwoong Yoon +1
Improving the reasoning capabilities of diffusion-based large language models (dLLMs) through reinforcement learning (RL) remains an open problem. The intractability of dLLMs likel…
Robust Adversarial Reinforcement Learning in Stochastic Games via Sequence Modeling
Xiaohang Tang, Zhuowen Cheng, Satyabrat Kumar
The Transformer, a highly expressive architecture for sequence modeling, has recently been adapted to solve sequential decision-making, most notably through the Decision Transforme…