From the 1 of 6 linked papers with an AI index.
6 papers
PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective
Shengtian Yang, Yewen Li, Peng Jiang +4
The paper introduces PlatformBid, a benchmark for evaluating auto-bidding algorithms from the perspective of a unified advertising platform that combines SSP, DSP, and ad exchange…
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
Kaibing Yang, Guangfeng Cai, Shengtian Yang +6
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Guangfeng Cai, Kaibing Yang, Shuo He +4
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…
Phase-Aware Mixture of Experts for Agentic Reinforcement Learning
Shengtian Yang, Yu Li, Shuo He +4
Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
Yu Li, Guangfeng Cai, Shengtian Yang +5
Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool pl…
MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection
Shengtian Yang, Yue Feng, Yingshi Liu +2
Video Anomaly Detection (VAD) aims to locate unusual activities or behaviors within videos. Recently, offline VAD has garnered substantial research attention, which has been invigo…