1 citations · 1 across the 5 of their papers we have counts for
8 papers
InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
Kun Luo, Hongjin Qian, Zheng Liu +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a promising approach for enhancing agentic deep search. However, its application is often hindered by low \textbf{Reward De…
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
Zhuoran Jin, Hongbang Yuan, Kejian Zhu +5
Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainl…
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
Tianyi Men, Zhuoran Jin, Pengfei Cao +3
As Multimodal Large Language Models (MLLMs) advance, multimodal agents show promise in real-world tasks like web navigation and embodied intelligence. However, due to limitations i…
ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework
Ao Chang, Tong Zhou, Yubo Chen +4
Legal Judgment Prediction (LJP) aims to predict judicial outcomes, including relevant legal charge, terms, and fines, which is a crucial process in Large Language Model(LLM). Howev…
RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
Chenlong Zhang, Zhuoran Jin, Hongbang Yuan +5
The widespread deployment of Large Language Models (LLMs) trained on massive, uncurated corpora has raised growing concerns about the inclusion of sensitive, copyrighted, or illega…
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
Ning Yang, Fangxin Liu, Junjie Wang +4
Large language models (LLMs) have achieved remarkable performance across a wide range of NLP tasks. However, their substantial inference cost poses a major barrier to real-world de…