6 citations · 9 across the 8 of their papers we have counts for
7 papers · 1 filter
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
Vedant Shah, Johan Obando-Ceron, Vineet Jain +10
The reasoning performance of large language models (LLMs) can be substantially improved by training them with reinforcement learning (RL). The RL objective for LLM training involve…
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
Siddarth Venkatraman, Vineet Jain, Sarthak Mittal +9
Test-time scaling methods improve the capabilities of large language models (LLMs) by increasing the amount of compute used during inference to make a prediction. Inference-time co…
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
Tingting Chen, Beibei Lin, Srinivas Anumasa +5
Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in noisy, open-ended scientific environments,…
Masked Generative Priors Improve World Models Sequence Modelling Capabilities
Cristian Meo, Mircea Lica, Zarif Ikram +6
Deep Reinforcement Learning (RL) has become the leading approach for creating artificial agents in complex environments. Model-based approaches, which are RL methods with world mod…
Towards DNA-Encoded Library Generation with GFlowNets
Michał Koziarski, Mohammed Abukalam, Vedant Shah +7
DNA-encoded libraries (DELs) are a powerful approach for rapidly screening large numbers of diverse compounds. One of the key challenges in using DELs is library design, which invo…
Efficient Causal Graph Discovery Using Large Language Models
Thomas Jiralerspong, Xiaoyin Chen, Yash More +2
We propose a novel framework that leverages LLMs for full causal graph discovery. While previous LLM-based methods have used a pairwise query approach, this requires a quadratic nu…