bilingual document corpus 1document VQA 1layout and table analysis 1multimodal document parsing 1multi-task reinforcement learning 1synthetic data generation 1
From the 1 of 20 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
Long Li, Zhijian Zhou, Jiaran Hao +9
A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pa…
cs.LG2025
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
Shuyao Xu, Cheng Peng, Jiangxuan Long +3
Recent advances in model distillation show that data from advanced reasoning models can effectively train smaller student models. However, standard practices discard incorrect reas…
cs.LG2024
Robust Deep Hawkes Process under Label Noise of Both Event and Occurrence
Xiaoyu Tan, Bin Li, Xihe Qiu +3
Integrating deep neural networks with the Hawkes process has significantly improved predictive capabilities in finance, health informatics, and information technology. Nevertheless…