From the 1 of 21 linked papers with an AI index.
21 papers
Meta-Learning Preferences for Multilingual LLM Alignment
Jiaying Lin, Seongho Son, Nam Phuong Tran +3
The paper introduces a meta-learning method that uses preference data from high-resource languages to quickly adapt large language models to low-resource languages with very few hu…
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks
Seongho Son, Sangwoong Yoon, Jiahua Tang +3
Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many iss…
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Juliusz Ziomek, William Bankes, Lorenz Wolf +3
We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, models must efficiently navigate Wi…
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
Stipe Frkovic, Metod Jazbec, Dan Zhang +3
Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel tok…
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham +3
Modern action-conditioned video world models achieve strong short-horizon visual realism, yet remain unreliable on rare, interaction-critical transitions that dominate downstream p…
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
Xiaohang Tang, Keyue Jiang, Che Liu +4
Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likeli…