4 papers
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
Wei-Chia Chang, Yan-Ann Chen
Vehicle make and model recognition (VMMR) is an important task in intelligent transportation systems, but existing approaches struggle to adapt to newly released models. Contrastiv…
From Faithfulness to Correctness: Generative Reward Models that Think Critically
Qiyao Ma, Yunsheng Shi, Hongtao Tian +3
Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathemat…
WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
Junyu Wu, Weiming Chang, Xiaotao Liu +10
Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent paradigm for training large language models and multimodal systems. Despite the notable advances enable…
G-Core: A Simple, Scalable and Balanced RLHF Trainer
Junyu Wu, Weiming Chang, Xiaotao Liu +8
Reinforcement Learning from Human Feedback (RLHF) has become an increasingly popular paradigm for training large language models (LLMs) and diffusion models. While existing RLHF tr…