collaborators

10 papers

cs.CR2026

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Yongli Xiang, Zhifang Zhang, Bojun Yang +4

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrat…

cs.LG2026

PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective

Shengtian Yang, Yewen Li, Peng Jiang +4

The paper introduces PlatformBid, a benchmark for evaluating auto-bidding algorithms from the perspective of a unified advertising platform that combines SSP, DSP, and ad exchange…

cs.LG2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Kaibing Yang, Guangfeng Cai, Shengtian Yang +6

Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…

cs.CL2026

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

Guangfeng Cai, Kaibing Yang, Shuo He +4

Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…

cs.AI2026

Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation

Jianan Chen, Zhifang Zhang, Shuo He +3

Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of…

cs.CV2026

Test-Time Attention Purification for Backdoored Large Vision Language Models

Zhifang Zhang, Bojun Yang, Shuo He +5

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…