104 citations · 144 across the 18 of their papers we have counts for
11 papers · 1 filter
Perception-Aware Policy Optimization for Multimodal Reasoning
Zhenhailong Wang, Xuehang Guo, Sofia Stoica +8
Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endowing Large Language Models (LLMs) with robust multi-step reasoning abilit…
Acting Less is Reasoning More! Teaching Model to Act Efficiently
Hongru Wang, Cheng Qian, Wanjun Zhong +7
Tool-integrated reasoning (TIR) augments large language models (LLMs) with the ability to invoke external tools during long-form reasoning, such as search engines and code interpre…
Graph Foundation Models: A Comprehensive Survey
Zehong Wang, Zheyuan Liu, Tianyi Ma +16
Graph-structured data pervades domains such as social networks, biological systems, knowledge graphs, and recommender systems. While foundation models have transformed natural lang…
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
Cheng Qian, Hongyi Du, Hongru Wang +6
Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity…
DecisionFlow: Advancing Large Language Model as Principled Decision Maker
Xiusi Chen, Shanyong Wang, Cheng Qian +3
In high-stakes domains such as healthcare and finance, effective decision-making demands not just accurate outcomes but transparent and explainable reasoning. However, current lang…
RM-R1: Reward Modeling as Reasoning
Xiusi Chen, Gaotang Li, Ziqi Wang +9
Reward modeling is essential for aligning large language models with human preferences through reinforcement learning. To provide accurate reward signals, a reward model (RM) shoul…