5 papers
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
Xia Zeng, Yihan Chen, Luhui Liu +3
We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The agent must follow a multi-stage St…
Advancing Open-source World Models
Robbyant Team, Zelin Gao, Qiuyu Wang +21
We present LingBot-World, an open-sourced world simulator stemming from video generation. Positioned as a top-tier world model, LingBot-World offers the following features. (1) It…
An Index-based Approach for Efficient and Effective Web Content Extraction
Yihan Chen, Benfeng Xu, Xiaorui Wang +1
As web agents (e.g., Deep Research) routinely consume massive volumes of web pages to gather and analyze information, LLM context management -- under large token budgets and low si…
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
Yihan Chen, Benfeng Xu, Xiaorui Wang +2
Autonomous agents, which perceive environments and take actions to achieve goals, have become increasingly feasible with the advancements in large language models (LLMs). However,…
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
Yihang Chen, Haikang Deng, Kaiqiao Han +1
Chain-of-Thought (CoT) reasoning enhances large language models (LLMs) by decomposing complex problems into step-by-step solutions, improving performance on reasoning tasks. Howeve…