5 papers
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
Xianpeng Zhang, Jiahua Yang, Dongyu Chen +7
The paper presents a benchmark for multimodal long-document summarization and a two-stage training framework (MMLDSum-LLM) that incorporates visual-alignment and keyword-aware loss…
expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the…
fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the…
Learning from Prompt itself: the Hierarchical Attribution Prompt Optimization
Dongyu Chen, Jian Ma, Xianpeng Zhang +5
Optimization is fundamental across numerous disciplines, typically following an iterative process of refining an initial solution to enhance performance. This principle is equally…
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
Zhiwei Jin, Xiaohui Song, Nan Wang +36
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hu…