7 papers
Insertion Based Sequence Generation with Learnable Order Dynamics
Dhruvesh Patel, Benjamin Rozonoyer, Gaurav Pandey +3
Existing insertion-based masked diffusion models that generate sequences by interleaving token insertion with unmasking use fixed schedules that are not dependent on the data. For…
Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo
While long, explicit chains-of-thought (CoT) have proven effective on complex reasoning tasks, they are costly to generate during inference. Non-verbal reasoning methods have emerg…
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Young-Suk Lee, Ramon Fernandez Astudillo, Radu Florian
Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in as…
Do LLMs Benefit From Their Own Words?
Jenny Y. Huang, Leshem Choshen, Wei Sun +6
In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses. We revisit this design choi…
Latent Principle Discovery for Language Model Self-Improvement
Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo
When language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. Howev…
Optimal Policy Minimum Bayesian Risk
Ramón Fernandez Astudillo, Md Arafat Sultan, Aashka Trivedi +4
Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniq…