7 citations · 23 across the 22 of their papers we have counts for
4 papers · 1 filter
SocraticPO: Policy Optimization via Interactive Guidance
Zirui Liu, Tingyue Pan, Jie Ouyang +8
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…
Deep Thinking by Markov Chain of Continuous Thoughts
Jiayu Liu, Zhenya Huang, Xuan Yang +6
Transformer-based models can perform complicated reasoning by generating reasoning paths token by token. While effective, this approach often requires generating thousands of token…
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
Tong Xiao, Xin Xu, Zhenya Huang +4
Enhancing the multimodal reasoning capabilities of Multimodal Large Language Models (MLLMs) is a challenging task that has attracted increasing attention in the community. Recently…
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
Yan Zhuang, Qi Liu, Haoyang Bi +12
Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual perfo…