collaborators

13 papers

cs.AI2026

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

Yijun Zhang, Yule Xie, Jiaxin Ding +4

Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabi…

cs.CV2026

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

Yongkang Zhou, Xiang Xia, Cheng Yan +2

Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial re…

cs.AI2026

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

Cheng Yan, Guangyang Ye, Wuyang Zhang +5

While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and…

cs.AI2026

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Yijun Zhang, Fan Xu, Jiaxin Ding +6

Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of interme…

cs.IR2026

Multimodal Representation Alignment for Cross-modal Information Retrieval

Fan Xu, Luis A. Leiva

Different machine learning models can represent the same underlying concept in different ways. This variability is particularly valuable for in-the-wild multimodal retrieval, where…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…