1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2025★ 1 cited
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
ByteDance Seed, :, Jiaze Chen +267
We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 8…
cs.AI2024
Robot Policy Learning with Temporal Optimal Transport Reward
Yuwei Fu, Haichao Zhang, Di Wu +2
Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle thi…
cs.LG2024★ 1 cited
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
Yuwei Fu, Haichao Zhang, Di Wu +2
In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pr…