2 citations · 2 across the 6 of their papers we have counts for
15 papers
FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification
Zheming Fu, Ruizhe He, Wei Shang +4
Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gr…
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
Chen He, Yuhao Wu, Lei Wang +2
Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcome…
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-intern…
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Test-time reinforcement learning (TTRL) reports substantial accuracy gains on mathematical reasoning benchmarks using majority vote as a pseudo-label signal. We argue these gains a…
Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment
Haozhe Jia, Pengyu Yin, Wenshuo Chen +6
Physics-informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and prone to shortcut learning under s…
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics
Zhaoxin Yu, Nan Xu, Kun Chen +3
With the continuous advancement of reasoning abilities in Large Language Models (LLMs), their application to scientific reasoning tasks has gained significant research attention. C…