3 papers
cs.CL2026
ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
Shi-Qi Yan, Chao-Hong Tan, Qian Chen +3
Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GR…
cs.CL2026
Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens
Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan +4
The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical…
cs.CL2025
RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation
Shi-Qi Yan, Quan Liu, Zhen-Hua Ling
While Retrieval-Augmented Generation (RAG) has exhibited promise in utilizing external knowledge, its generation process heavily depends on the quality and accuracy of the retrieve…