3 papers
cs.LG2025
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
Feng Chen, Allan Raventos, Nan Cheng +2
Recent progress in large language models (LLMs) highlights the power of scaling test-time compute to achieve strong performance on complex tasks, such as mathematical reasoning and…
cs.LG2024
Is AI Robust Enough for Scientific Research?
Jun-Jie Zhang, Jiahao Song, Xiu-Cheng Wang +14
We uncover a phenomenon largely overlooked by the scientific community utilizing AI: neural networks exhibit high susceptibility to minute perturbations, resulting in significant d…
cs.LG2024
Symmetry Breaking in Neural Network Optimization: Insights from Input Dimension Expansion
Jun-Jie Zhang, Nan Cheng, Fu-Peng Li +4
Understanding the mechanisms behind neural network optimization is crucial for improving network design and performance. While various optimization techniques have been developed,…