The model is a component. The system around it is the product.
Every AI product is 20% model and 80% infrastructure. Data pipeline, retrieval system, evaluation harness, guardrails, latency optimisation, cost controls. Most teams discover this after shipping the demo — when users report confident wrong answers, hallucinations about private data and 8-second response times. PROPELOO designs the full AI system before writing the first prompt. The model is selected last.
Frequently Asked Questions
When should we use RAG vs fine-tuning?
Use RAG when your knowledge base changes frequently, when you need source citations, or when you cannot afford to retrain for every knowledge update. Use fine-tuning when you need to change the model's output format, style or tone rather than its knowledge — for example, teaching it to respond in your brand voice or produce a specific JSON structure. In most production systems, RAG handles dynamic knowledge and fine-tuning handles style/format. The combination is more powerful than either alone.
How do we evaluate whether the AI is actually working?
You need an evaluation pipeline that runs continuously, not a manual spot-check before each release. Define success metrics before building — task completion rate, factual accuracy against ground truth, hallucination rate, latency p50/p95. Run automated evaluations using a test set of representative queries with expected outputs. Use LLM-as-judge for open-ended quality assessment at scale. Track these metrics over time, especially after model updates or prompt changes. Without this, you cannot detect quality degradation.
How do you prevent prompt injection attacks?
Prompt injection prevention requires multiple layers: input validation and sanitisation before injection into prompts, clear separation between system instructions and user input using structured message formats, output filtering that detects instruction-following deviations, sandboxed tool execution that limits what agents can do even if hijacked, and monitoring for anomalous output patterns. Do not rely solely on the model refusing injections — LLMs are not reliable security boundaries.
What is the difference between an AI agent and a chatbot?
A chatbot responds to single-turn queries within a conversation. An AI agent executes multi-step tasks, uses tools (APIs, databases, browsers, code execution), maintains state across steps, makes decisions about which steps to take next and can complete tasks that take minutes to hours. Agents require significantly more engineering — task decomposition logic, tool definitions, error handling, retry strategies and human-in-the-loop gates for high-stakes actions. Most production 'agents' are actually sophisticated chatbots with tool calling.
How do we control AI infrastructure costs?
Cost control is an architectural design problem. Key levers: model tier routing (use GPT-4o-mini for simple queries, GPT-4o for complex), semantic caching (return cached responses for semantically similar queries), context window management (don't send unnecessary context), batch inference for non-real-time tasks, and retrieval precision tuning (retrieve less but better context). Well-designed cost controls typically reduce AI inference cost by 40-70% vs naive implementations while maintaining or improving quality.
Can we run AI models on-premise to keep data private?
Yes — models like Llama 3.1, Mistral Large, Mixtral 8x7B and Phi-3 can be self-hosted on GPU infrastructure. AWS Bedrock and Azure OpenAI offer managed APIs where data stays within your cloud environment without leaving to third-party model providers. Self-hosting requires GPU instance management, model serving infrastructure (vLLM, TGI, Ollama) and ongoing model update management. The performance and capability gap between self-hosted open models and frontier models (GPT-4o, Claude) is closing but not zero.
How do you handle PII in AI systems?
PII handling in AI systems requires: identification and redaction of PII before ingestion into RAG corpora or fine-tuning datasets, per-user access control at the retrieval layer to prevent cross-user data leakage, data retention policies that account for vector embeddings (deleting a source document does not automatically delete its embeddings), and output monitoring for PII appearing in model responses. GDPR right-to-deletion creates an obligation to remove not just source data but all derived vector representations.
What latency can we expect from a production AI system?
Time-to-first-token (TTFT) for streaming responses from OpenAI GPT-4o is typically 300-800ms. Full response latency for a RAG query (retrieval + generation) ranges from 1.5-8 seconds depending on retrieval complexity, context length and model. Semantic caching reduces latency to <50ms for cached queries. If your product requires sub-1-second full responses, this constrains model selection, context length and retrieval strategy significantly. We design latency budgets at the architecture stage.
How do you prevent AI hallucinations?
Hallucinations cannot be eliminated — they are a fundamental property of LLMs. They can be reduced through: retrieval augmentation that grounds responses in source documents, requiring the model to cite specific passages, output validation against structured schemas, confidence scoring and uncertainty quantification, restricting the model's response scope to retrieved context, and evaluation pipelines that measure hallucination rate continuously. For high-stakes decisions, human review queues for low-confidence outputs are the last line of defence.
How long does a production AI system take to build?
A well-scoped RAG system over a private document corpus — ingestion pipeline, retrieval, LLM integration, evaluation framework, frontend interface — takes 8-14 weeks from architecture to production. A multi-tool AI agent with custom integrations takes 12-20 weeks. Fine-tuned model deployment adds 4-6 weeks for dataset preparation, training and evaluation. The evaluation framework and monitoring infrastructure typically take as long as the core system — teams that skip this ship AI systems they cannot measure or maintain.