3 papers
cs.CV2025
Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
Yana Wei, Zeen Chi, Chongyu Wang +4
In open-world environments, human-object interactions (HOIs) evolve continuously, challenging conventional closed-world HOI detection models. Inspired by humans' ability to progres…
cs.CL2025
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
ByteDance Seed, :, Jiaze Chen +267
We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 8…
cs.AI2024
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
Ruijie Xu, Zhihan Liu, Yongfei Liu +4
We address the challenge of online Reinforcement Learning from Human Feedback (RLHF) with a focus on self-rewarding alignment methods. In online RLHF, obtaining feedback requires i…