3 papers
cs.CL2026
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
Yanlin Lai, Mitt Huang, Hangyu Guo +11
Reinforcement Learning from Human Feedback (RLHF) remains indispensable for aligning large language models (LLMs) in subjective domains. To enhance robustness, recent work shifts t…
cs.CV2026
STEP3-VL-10B Technical Report
Ailin Huang, Chengyuan Yao, Chunrui Han +90
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…
cs.SE2025
SolSearch: An LLM-Driven Framework for Efficient SAT-Solving Code Generation
Junjie Sheng, Yanqiu Lin, Jiehao Wu +4
The Satisfiability (SAT) problem is a core challenge with significant applications in software engineering, including automated testing, configuration management, and program verif…