collaborators

5 papers

cs.CV2025

Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs

Rujiao Long, Yang Li, Xingyao Zhang +7

Exploration capacity shapes both inference-time performance and reinforcement learning (RL) training for large (vision-) language models, as stochastic sampling often yields redund…

cs.AI2025

RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models

Tianqianjin Lin, Xi Zhao, Xingyao Zhang +5

Reinforcement learning (RL) can refine the reasoning abilities of large language models (LLMs), but critically depends on a key prerequisite: the LLM can already generate high-util…

cs.LG2025

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Weixun Wang, Shaopan Xiong, Gengru Chen +38

We introduce ROLL, an efficient, scalable, and user-friendly library designed for Reinforcement Learning Optimization for Large-scale Learning. ROLL caters to three primary user gr…

cs.AI2025

Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation

Yijia Luo, Yulin Song, Xingyao Zhang +5

Recent advancements in large language models (LLMs) have demonstrated remarkable reasoning capabilities through long chain-of-thought (CoT) reasoning. The R1 distillation scheme ha…

cs.CL2025

ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models

Haibin Chen, Kangtao Lv, Chengwei Hu +8

With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilit…