Reasoning Beyond Limits: Advances and Open Problems for LLMs
arXiv:2503.22732 · doi:10.1016/j.icte.2025.09.003
Abstract
Recent breakthroughs in generative reasoning have fundamentally reshaped how large language models (LLMs) address complex tasks, enabling them to dynamically retrieve, refine, and organize information into coherent multi-step reasoning chains. Techniques such as inference-time scaling, reinforcement learning, supervised fine-tuning, and distillation have been effectively applied to state-of-the-art models, including DeepSeek-R1, OpenAI o1 and o3, GPT-4o, Qwen-32B, and various Llama variants, significantly enhancing their reasoning capabilities. In this paper, we present a comprehensive review of the top 27 LLMs released between 2023 and 2025, such as Mistral AI Small 3 24B, DeepSeek-R1, Search-o1, QwQ-32B, and Phi-4, and analyze their core innovations and performance improvements. We also provide a detailed overview of recent advancements in multilingual large language models (MLLMs), emphasizing methods that improve cross-lingual reasoning and address the limitations of English-centric training. In parallel, we present a comprehensive review of progress in state space model (SSM)-based architectures, including models such as Mamba, which demonstrate improved efficiency for long-context processing compared to transformer-based approaches. Our analysis covers training strategies including general optimization techniques, mixture-of-experts (MoE) configurations, retrieval-augmented generation (RAG), chain-of-thought prompting, self-improvement methods, and test-time compute scaling and distillation frameworks. Finally, we identify key challenges for future research, including enabling multi-step reasoning without human supervision, improving robustness in chained task execution, balancing structured prompting with generative flexibility, and enhancing the integration of long-context retrieval and external tools.
The paper is published ICT Express Volume 11, Issue 6, December 2025, Pages 1054-1096
References in corpus (96)
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Scaling Instruction-Finetuned Language Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- A Survey on Multimodal Large Language Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Reflexion: Language Agents with Verbal Reinforcement Learning
- DeepSeek-V3 Technical Report
- Self-Refine: Iterative Refinement with Self-Feedback
- STaR: Bootstrapping Reasoning With Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Qwen2.5-VL Technical Report
- Cognitive Architectures for Language Agents
- Jamba: A Hybrid Transformer-Mamba Language Model
- InternLM2 Technical Report
- SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- From Sparse to Soft Mixtures of Experts
- Mixture-of-Agents Enhances Large Language Model Capabilities
- MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
- KTO: Model Alignment as Prospect Theoretic Optimization
- Solving math word problems with process- and outcome-based feedback
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
- State Space Model for New-Generation Network Alternative to Transformers: A Survey
- Training Large Language Models to Reason in a Continuous Latent Space
- LoRA Learns Less and Forgets Less
- Agentless: Demystifying LLM-based Software Engineering Agents
- Multi-Step Reasoning with Large Language Models, a Survey
- Pre-Trained Language Models for Keyphrase Prediction: A Review
- Chain-of-Thought Reasoning Without Prompting
- s1: Simple test-time scaling
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Mamba4Rec: Towards Efficient Sequential Recommendation with Selective State Space Models
- BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents
- Moshi: a speech-text foundation model for real-time dialogue
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
- Training Language Models to Self-Correct via Reinforcement Learning
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
- Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
- Reinforcement Learning Enhanced LLMs: A Survey
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
- TransMLA: Multi-Head Latent Attention Is All You Need
- UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- YAYI 2: Multilingual Open-Source Large Language Models
- Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
- V-STaR: Training Verifiers for Self-Taught Reasoners
- Understanding the performance gap between online and offline alignment algorithms
- Self-Taught Evaluators
- LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
- Contextual Document Embeddings
- KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
- A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
- Thinking LLMs: General Instruction Following with Thought Generation
- Stream of Search (SoS): Learning to Search in Language
- Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
- The Perfect Blend: Redefining RLHF with Mixture of Judges
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
- Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
- TigerBot: An Open Multilingual Multitask LLM
- Pyramidal Flow Matching for Efficient Video Generative Modeling
- HMoE: Heterogeneous Mixture of Experts for Language Modeling
- Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
- Evolving Deeper LLM Thinking
- Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
- Following Length Constraints in Instructions
- Process Reinforcement through Implicit Rewards
- Orion-14B: Open-source Multilingual Large Language Models
- Not All LLM Reasoners Are Created Equal
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- Monolingual or Multilingual Instruction Tuning: Which Makes a Better Alpaca
- BOND: Aligning LLMs with Best-of-N Distillation
- Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
- ReMamba: Equip Mamba with Effective Long-Sequence Modeling
- Free Process Rewards without Process Labels
- Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
- Direct Judgement Preference Optimization
- Critique-out-Loud Reward Models
- SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling