8 citations · 18 across the 13 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
First Return, Entropy-Eliciting Explore
Tianyu Zheng, Tianshun Xing, Qingshui Gu +10
Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (…
cs.AI2025
Aligning Instruction Tuning with Pre-training
Yiming Liang, Tianyu Zheng, Xinrun Du +12
Instruction tuning enhances large language models (LLMs) to follow human instructions across diverse tasks, relying on high-quality datasets to guide behavior. However, these datas…
cs.AI2024★ 1 cited
MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces
Tianyu Zheng, Ge Zhang, Xingwei Qu +3
Drawing upon the intuition that aligning different modalities to the same semantic embedding space would allow models to understand states and actions more easily, we propose a new…