activity
20242026
collaborators

18 papers

cs.LG2026

SocraticPO: Policy Optimization via Interactive Guidance

Zirui Liu, Jie Ouyang, Qi Liu +8

Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…

cs.CY2026

Agent4Edu: Generating Learner Response Data by Generative Agents for Intelligent Education Systems

Weibo Gao, Qi Liu, Linan Yue +5

Personalized learning represents a promising educational strategy within intelligent educational systems, aiming to enhance learners' practice efficiency. However, the discrepancy…

cs.LG2026

Deep Thinking by Markov Chain of Continuous Thoughts

Jiayu Liu, Zhenya Huang, Xuan Yang +6

Transformer-based models can perform complicated reasoning by generating reasoning paths token by token. While effective, this approach often requires generating thousands of token…

cs.LG2026

Survey of Computerized Adaptive Testing: A Machine Learning Perspective

Yan Zhuang, Qi Liu, Haoyang Bi +12

Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual perfo…

cs.LG2026

Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward

Tong Xiao, Xin Xu, Zhenya Huang +4

Enhancing the multimodal reasoning capabilities of Multimodal Large Language Models (MLLMs) is a challenging task that has attracted increasing attention in the community. Recently…

cs.IR2025

A Survey on Deep Text Hashing: Efficient Semantic Text Retrieval with Binary Representation

Liyang He, Zhenya Huang, Cheng Yang +6

With the rapid growth of textual content on the Internet, efficient large-scale semantic text retrieval has garnered increasing attention from both academia and industry. Text hash…