5 papers
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
Chen Li, Nazhou Liu, Kai Yang
Since DeepSeek-R1 popularized, Group Relative Policy Optimization (GRPO) has become the core part of training Reasoning LLMs. However, we find some deficiency that influences RL st…
Contributions to Robust and Efficient Methods for Analysis of High Dimensional Data
Kai Yang
A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is…
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
Junjie Gao, Runze Liu, Yingzhe Peng +4
Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide…
News Source Citing Patterns in AI Search Systems
Kai-Cheng Yang
AI-powered search systems are emerging as new information gatekeepers, fundamentally transforming how users access news and information. Despite their growing influence, the citati…
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Yingzhe Peng, Gongrui Zhang, Miaosen Zhang +7
Enhancing reasoning in Large Multimodal Models (LMMs) faces unique challenges from the complex interplay between visual perception and logical reasoning, particularly in compact 3B…