Frequently Asked Questions
What is RAG and why do you recommend it over fine-tuning?
RAG (Retrieval-Augmented Generation) retrieves relevant documents from your knowledge base at query time and includes them in the LLM prompt — grounding the response in verified content rather than the model's training data. Fine-tuning trains the model on your data, which changes the model's weights but does not prevent hallucination, cannot be updated without retraining, and costs thousands of dollars per run. RAG is cheaper, instantly updatable with new content, and produces traceable, source-cited responses. We recommend fine-tuning only for specific style/format requirements that cannot be achieved through prompting.
Which LLM should we use?
For most production chatbots: GPT-4o or Claude 3.5 Sonnet. They have the best instruction following, lowest hallucination rates and strongest reasoning among commercially available models. For document-heavy RAG with long context: Claude 3.5 Sonnet's 200K context window is an advantage. For data privacy requirements where data cannot leave your infrastructure: Llama 3 70B or Mistral self-hosted on AWS/GCP. For high-volume, cost-sensitive applications: GPT-4o-mini or Claude 3 Haiku with RAG grounding performs well at a fraction of the cost.
How do you prevent hallucinations?
RAG grounding is the primary mechanism — by constraining the LLM to answer from retrieved documents, you eliminate the model's tendency to invent information. Additional measures: confidence thresholds that return "I don't know" rather than a low-confidence answer, source citation for every factual claim, output validation against the retrieved documents, and human review queues for responses flagged as low-confidence. No system eliminates hallucination entirely — monitoring and feedback loops are required.
How much does it cost to run in production?
Costs depend on query volume and LLM choice. GPT-4o at $5/1M input tokens: a 1,000-token query (prompt + context) costs $0.005. At 10,000 queries/day: $50/day in LLM costs. Claude 3.5 Sonnet is similar pricing. Vector database costs (Pinecone): $70/month for 1M vectors. Embedding generation is approximately 10x cheaper than inference. For high-volume applications, GPT-4o-mini or self-hosted open-source models can reduce LLM costs by 10-20x with acceptable quality.
Can the chatbot integrate with our CRM / support ticketing system?
Yes. LLM tool use (function calling) allows the chatbot to query your CRM, create support tickets, look up order status, check account information and trigger workflows — all within the conversation. We build tool definitions for your specific integrations, handle authentication securely, and implement appropriate data access controls so the chatbot only accesses information the user is authorised to see.
How do you handle multi-language support?
Modern LLMs handle multi-language input and output natively without any additional configuration — GPT-4o and Claude 3.5 Sonnet both support 50+ languages. The RAG layer requires language-appropriate chunking and embedding models that perform well in your target languages (Cohere embed-v3 has strong multilingual performance). We recommend testing retrieval quality in each target language with a representative question set before launch.
What is the typical development timeline?
A production-ready RAG chatbot with a defined knowledge base: 4–8 weeks. This includes document ingestion pipeline, vector database setup, prompt engineering, guardrails, evaluation framework and API development. Integrations with existing systems (CRM, ticketing, calendar) add 1–2 weeks each. Voice interfaces add 2–3 weeks. Custom fine-tuning adds 3–6 weeks depending on dataset preparation requirements.
How do we measure if the chatbot is working?
The key metrics are: retrieval precision (percentage of retrieved chunks relevant to the query), answer faithfulness (percentage of response content grounded in retrieved documents), answer relevance (percentage of responses that actually answer the question asked), and user satisfaction (thumbs up/down, escalation rate to human agents). RAGAS provides automated measurement of the first three. User feedback and human evaluation cover the fourth. We set up these measurement systems before launch, not after.