collaborators

10 papers

cs.CL2026

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

Zihao Deng, Yining Zhu, Leiming Wang +6

Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectorie…

cs.CL2026

FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards

Zihao Deng, Yining Zhu, Leiming Wang +6

Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume…

cs.LG2026

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

Weixin Wang, Yu Yang, Wei Deng +1

We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its weights. Recent Sequential Mon…

cs.LG2026

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

Yu Yang, Yihong Guo, Anqi Liu +1

Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under mismatched transition dynamics. Exi…

cs.CV2026

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

Wencan Jiang, Jiangning Zhang, Jianbiao Mei +6

Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Existing approaches often trade…

cs.LG2026

MOBODY: Model Based Off-Dynamics Offline Reinforcement Learning

Yihong Guo, Yu Yang, Pan Xu +1

We study off-dynamics offline reinforcement learning, where the goal is to learn a policy from offline source and limited target datasets with mismatched dynamics. Existing methods…