30 papers
L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
Boqin Yuan, Ziqi Wang
The paper introduces L-MARS, a multi‑agent system for legal question answering that uses agentic search and a judge‑driven loop to verify that citations actually support each claim…
Kwai Summary Attention Technical Report
Chenglong Chu, Guorui Zhou, Guowang Zhang +35
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agen…
OneReason Technical Report
OneRec Team, Biao Yang, Boyang Ding +81
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. Howev…
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
Ziqi Wang, Chang Che, Qi Wang +4
While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safe…
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
Minxuan Lv, Tiehua Mei, Tanlong Du +9
We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Existing long-context RL…