From the 1 of 8 linked papers with an AI index.
8 papers
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
Kai Yang, Jingwei Xu, Wanyu Wang +4
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradat…
Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update
Daocheng Fu, Rong Wu, Yu Yang +7
The paper introduces PUST, a framework that uses a lightweight proxy model to explore high‑reward behaviors and then transfers the relative improvement signals to a larger primary…
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
Yang Chen, Yufan Shen, Wenxuan Huang +7
Multimodal Large Language Models (MLLMs) exhibit impressive performance across various visual tasks. Subsequent investigations into enhancing their visual reasoning abilities have…
O-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
Jianbiao Mei, Tao Hu, Daocheng Fu +11
Large Language Models (LLMs), despite their advancements, are fundamentally limited by their static parametric knowledge, hindering performance on tasks requiring open-domain up-to…
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
Siqi Li, Yufan Shen, Xiangnan Chen +13
The rapid advancement of multimodal large language models (MLLMs) has profoundly impacted the document domain, creating a wide array of application scenarios. This progress highlig…
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
Renqiu Xia, Bo Zhang, Hancheng Ye +9
Recently, many versatile Multi-modal Large Language Models (MLLMs) have emerged continuously. However, their capacity to query information depicted in visual charts and engage in r…