5 papers
Multi-Branch Policy Optimization for Multimodal Large Language Models
Shuai Lyu, Yuning Gong, Ruiling Gao +7
Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens i…
Scaling Reinforcement Learning for Content Moderation with Large Language Models
Hamed Firooz, Rui Liu, Yuchen Lu +15
Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evalua…
Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
Yong Deng, Guoqing Wang, Zhenzhe Ying +12
Large language models (LLMs) exhibit remarkable problem-solving abilities, but struggle with complex tasks due to static internal knowledge. Retrieval-Augmented Generation (RAG) en…
SQL-o1: A Self-Reward Heuristic Dynamic Search Method for Text-to-SQL
Shuai Lyu, Haoran Luo, Ripeng Li +6
Text-to-SQL (Text2SQL) aims to map natural language questions to executable SQL queries. Although large language models (LLMs) have driven significant progress, current approaches…
HGTUL: A Hypergraph-based Model For Trajectory User Linking
Fengjie Chang, Xinning Zhu, Zheng Hu +1
Trajectory User Linking (TUL), which links anonymous trajectories with users who generate them, plays a crucial role in modeling human mobility. Despite significant advancements in…