7 papers
HSAP: A Hierarchical Sequence-aware Parallelism for Hybrid-Context Generative Models
Songxin Zhang, Zejian Xie, Zhuoyang Song +4
In this paper, we aim to combine the advantages of existing sequence parallelism paradigms and overcomes their drawbacks, the most serious of which is the incapability to correctly…
The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment
Junyu Lu, Qi Wei, Peishuo Zheng +6
Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial r…
ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring
Deyi Ji, Junyu Lu, Xuanyi Liu +7
Online advertising governance faces significant challenges due to the non-stationary nature of regulatory policies, where emerging mandates (e.g., restrictions on education or aest…
SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
Wanxin Tian, Shijie Zhang, Kevin Zhang +12
Self-evolution, the ability of agents to autonomously improve their reasoning and behavior, is essential for the embodied domain with long-horizon, real-world tasks. Despite curren…
Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
Junyu Lu, Songxin Zhang, Zejian Xie +2
Recent advances in GUI agents have achieved remarkable grounding and action-prediction performance, yet existing models struggle with unreliable reward signals and limited online t…
L0: Reinforcement Learning to Become General Agents
Junjie Zhang, Jingyi Xi, Zhuoyang Song +7
Training large language models (LLMs) to act as autonomous agents for multi-turn, long-horizon tasks remains significant challenges in scalability and training efficiency. To addre…