What is the difference between RAG and fine-tuning?
RAG retrieves relevant documents at query time and grounds the LLM response in verified content — the knowledge is external and instantly updatable. Fine-tuning trains the model weights on your data — the knowledge is internal and cannot be updated without retraining. RAG prevents hallucination by constraining responses to retrieved content. Fine-tuning does not — a fine-tuned model still hallucinates. For knowledge base Q&A, RAG is almost always the correct choice. Fine-tuning is appropriate for response style, format or domain-specific reasoning patterns that cannot be achieved through prompting.