collaborators

5 papers

cs.AI2026

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

Yipeng Shi, Zhipeng Ma, Yue Wang +4

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimizati…

cs.AI2026

FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models

Kehan Jiang, Haonan Dong, Zhaolu Kang +2

Recent Large Reasoning Models (LRMs) like DeepSeek-R1 have demonstrated remarkable success in complex reasoning tasks, exhibiting human-like patterns in exploring multiple alternat…

cs.CV2025

LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts

Qifeng Cai, Hao Liang, Zhaoyang Han +6

Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from lim…

cs.CL2025

Let's Verify Math Questions Step by Step

Chengyu Shen, Zhen Hao Wong, Runming He +8

Large Language Models (LLMs) have recently achieved remarkable progress in mathematical reasoning. To enable such capabilities, many existing works distill strong reasoning models…

cs.CL2025

MathDebugger: Detecting and Diagnosing Errors in Synthetic Mathematical Data

Hao Liang, Meiyi Qiang, Yuying Li +7

Synthetic mathematical data has become an important resource for scaling the reasoning capabilities of large language models, yet errors in generated questions and solutions can su…