Guided Decoding and Its Critical Role in Retrieval-Augmented Generation
arXiv:2509.06631 · doi:10.1109/SIU66497.2025.11111950
Abstract
The integration of Large Language Models (LLMs) into various applications has driven the need for structured and reliable responses. A key challenge in Retrieval-Augmented Generation (RAG) systems is ensuring that outputs align with expected formats while minimizing hallucinations. This study examines the role of guided decoding in RAG systems, comparing three methods, Outlines, XGrammar, and LM Format Enforcer, across different multi-turn prompting setups (0-turn, 1-turn, and 2-turn). By evaluating success rates, hallucination rates, and output quality, we provide insights into their performance and applicability. Our findings reveal how multi-turn interactions influence guided decoding, uncovering unexpected performance variations that can inform method selection for specific use cases. This work advances the understanding of structured output generation in RAG systems, offering both theoretical insights and practical guidance for LLM deployment.
References in corpus (7)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- "We Need Structured Output": Towards User-centered Constraints on Large Language Model Output
- Efficient Guided Generation for Large Language Models
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- Guiding Language Models of Code with Global Context using Monitors
- Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs
- Reward-Guided Speculative Decoding for Efficient LLM Reasoning