Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Dynamic Chain-of-Thought: Towards Adaptive Deep Reasoning
Libo Wang
To reduce the cost and consumption of computing resources caused by computational redundancy and delayed reward assignment in long CoT, this research proposes the dynamic chain-of-…
cs.AI2025
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
Libo Wang
To address the sycophancy problem caused by reinforcement learning from human feedback in large language models, this research applies synthetic data intervention technology to the…