6 citations · 11 across the 4 of their papers we have counts for
9 papers
Reinforcement Learning Teachers of Test Time Scaling
Edoardo Cetin, Tianyu Zhao, Yujin Tang
Training reasoning language models (LMs) with reinforcement learning (RL) for one-hot correctness inherently relies on the LM being able to explore and solve its task with some cha…
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
Jeffrey Seely, Yuki Imajuku, Tianyu Zhao +2
Existing reasoning benchmarks for large language models (LLMs) frequently fail to capture authentic creativity, often rewarding memorization of previously observed patterns. We add…
Large Language Models to Diffusion Finetuning
Edoardo Cetin, Tianyu Zhao, Yujin Tang
We propose a new finetuning method to provide pre-trained large language models (LMs) the ability to scale test-time compute through the diffusion framework. By increasing the numb…
An Evolved Universal Transformer Memory
Edoardo Cetin, Qi Sun, Tianyu Zhao +1
Prior methods propose to offset the escalating costs of modern foundation models by dropping specific parts of their contexts with hand-designed rules, while attempting to preserve…
Multi-Referenced Training for Dialogue Response Generation
Tianyu Zhao, Tatsuya Kawahara
In open-domain dialogue response generation, a dialogue context can be continued with diverse responses, and the dialogue models should capture such one-to-many relations. In this…
Designing Precise and Robust Dialogue Response Evaluators
Tianyu Zhao, Divesh Lala, Tatsuya Kawahara
Automatic dialogue response evaluator has been proposed as an alternative to automated metrics and human evaluation. However, existing automatic evaluators achieve only moderate co…