1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Haotian Xu, Xing Wu, Weinong Wang +11
Can scaling transform reasoning? In this work, we explore the untapped potential of scaling Long Chain-of-Thought (Long-CoT) data to 1000k samples, pioneering the development of a…
cs.CL2025
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
Hantao Lou, Jiaming Ji, Kaile Wang +1
The rapid advancement of large language models (LLMs) has led to significant improvements in their capabilities, but also to increased concerns about their alignment with human val…
cs.AI2024
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Jiaming Ji, Jiayi Zhou, Hantao Lou +16
Reinforcement learning from human feedback (RLHF) has proven effective in enhancing the instruction-following capabilities of large language models; however, it remains underexplor…