PROPELOO

AI / INTELLIGENT SYSTEMS

Design AI systems that do more than generate answers.

PROPELOO engineers production AI systems — from RAG pipelines and LLM integration to autonomous agents, fine-tuning workflows and the evaluation infrastructure that makes AI outputs measurable. The model is a component. The system around it is the product.

The model is a component. The system around it is the product.

Every AI product is 20% model and 80% infrastructure. Data pipeline, retrieval system, evaluation harness, guardrails, latency optimisation, cost controls. Most teams discover this after shipping the demo — when users report confident wrong answers, hallucinations about private data and 8-second response times. PROPELOO designs the full AI system before writing the first prompt. The model is selected last.

What sits underneath a production AI system.

An LLM API call is not an AI product. It is a starting point. The retrieval, evaluation, orchestration and monitoring layers determine whether the product works.

System Layers

  • Data Ingestion & Processing: Document parsing, chunking strategy, cleaning, embedding pipeline, data refresh cadence, format normalisation
  • Vector Store & Retrieval: Vector database, retrieval strategy, hybrid search, re-ranking, context assembly, metadata filtering
  • Model Orchestration Layer: LLM selection, prompt engineering, model routing, structured output, function calling, fallback logic
  • Agent & Tool Layer: Tool definitions, multi-agent coordination, memory management, task decomposition, guardrails, escalation paths
  • Evaluation & Monitoring: Accuracy metrics, hallucination detection, latency tracking, cost monitoring, A/B evaluation, LLM-as-judge

Core Technical Capabilities

  • RAG Pipeline Architecture

    End-to-end retrieval-augmented generation pipelines — document ingestion, chunking strategy, embedding model selection, vector store configuration, hybrid search and re-ranking for production-quality retrieval.

  • LLM Integration & Routing

    Multi-model integration with intelligent routing — primary model, fallback, cost-optimised routing based on task complexity, and provider failover with latency-aware load balancing.

  • AI Agent Development

    Autonomous AI agents with tool use, multi-step task execution, memory systems, error handling, human-in-the-loop escalation paths and production-grade reliability engineering.

  • Fine-tuning & Alignment

    Domain-specific fine-tuning workflows — dataset preparation, PEFT/LoRA training, alignment techniques, evaluation against base model and deployment of fine-tuned models to production inference infrastructure.

  • Evaluation & Testing Framework

    Automated evaluation pipelines measuring accuracy, hallucination rate, task completion, latency and cost — including LLM-as-judge, human evaluation sampling and regression test suites for each model update.

  • AI Infrastructure Cost Control

    Token budget management, semantic caching, model tier routing, batch inference for non-real-time workloads and infrastructure rightsizing — reducing AI operational cost by 40–70% without degrading quality.

How we think about AI system design.

The model isn't the product. The system that feeds it context, constrains its outputs and measures its accuracy is.

  • Context is the bottleneck, not compute

    An LLM is only as useful as the context it receives. Chunking strategy, embedding model, retrieval configuration and re-ranking logic determine output quality more than model selection. We optimise retrieval before optimising models.

    Axiom:

  • Evaluation is not optional

    AI systems without evaluation pipelines degrade invisibly. Hallucination rate, task completion, factual accuracy and latency must be measured continuously — not manually spot-checked before each release.

    Axiom:

  • Agents are systems, not features

    An AI agent is a distributed system with tool dependencies, state management and failure modes. Production agents require the same reliability engineering as any microservice — retry logic, circuit breakers, fallback paths and monitoring.

    Axiom:

  • Cost is an architectural constraint

    A system that costs $0.40 per user interaction has a ceiling on its addressable market. Token efficiency, semantic caching, model tier routing and batch inference are architectural decisions made at design time.

    Axiom:

The decisions that define an AI system.

Model selection is the last decision, not the first. These choices come before it.

  • RAG vs fine-tuning vs both

    Impact: Fine-tuning bakes knowledge into weights that cannot be updated without retraining. RAG externalises knowledge. Choose based on how frequently the knowledge base changes.

    • RAG only — retrieval-augmented generation, no training cost, knowledge updateable without retraining
    • Fine-tuning only — task-specific behaviour, knowledge baked into weights, requires retraining for updates
    • RAG + fine-tuning — format/style adaptation with fine-tuning, dynamic knowledge via RAG
    • Prompt engineering only — fastest to start, limited control over output format and style
  • Single model vs model router

    Impact: A model router can reduce inference cost by 60-80% by using cheaper models for simple tasks. The routing logic itself introduces an additional failure mode.

    • Single model (GPT-4o, Claude Sonnet) — simple, consistent behaviour, single provider dependency
    • Multi-model router — route by task complexity, cost-optimise with GPT-4o-mini for simple tasks
    • Specialised models per domain — best quality per task, higher integration complexity
  • Synchronous vs streaming responses

    Impact: Streaming changes the frontend architecture and requires streaming-aware infrastructure at every layer. Non-streaming AI products feel slow regardless of latency.

    • Synchronous — simpler integration, poor UX for long responses (8-30 second waits)
    • Streaming (SSE/WebSocket) — real-time token-by-token output, better UX, more complex infrastructure
    • Async with polling — background processing, results fetched when ready, suitable for batch tasks
  • Managed API vs self-hosted model

    Impact: Data residency requirements and compliance posture drive this decision. Self-hosted models at scale require GPU infrastructure expertise.

    • OpenAI / Anthropic / Gemini API — fastest integration, data leaves your environment, per-token cost
    • AWS Bedrock / Azure OpenAI — managed API, data stays in your cloud, compliance-friendly
    • Self-hosted (Llama 3 / Mistral / Mixtral) — full data control, higher ops cost, compliance advantage
  • Vector database selection

    Impact: Vector store choice affects retrieval latency, metadata filtering capability and operational burden. Changing vector stores requires re-embedding the entire corpus.

    • Pinecone — fully managed, fast, no infrastructure overhead, per-vector cost
    • pgvector (PostgreSQL extension) — existing infra reuse, SQL joins on metadata, lower scale ceiling
    • Weaviate / Qdrant — self-hosted, full control, operational overhead
    • OpenSearch k-NN — AWS-native, good if already using OpenSearch
  • Agent framework vs custom orchestration

    Impact: Framework abstractions accelerate prototyping but introduce debugging complexity and version dependency risk. Production agents at scale often outgrow framework abstractions.

    • LangChain — extensive integrations, heavy abstraction layer, version instability
    • LlamaIndex — best for RAG-heavy systems, document-centric architecture
    • AutoGen / CrewAI — multi-agent frameworks, higher-level abstractions
    • Custom orchestration — full control, no abstraction overhead, higher build cost

What AI infrastructure becomes.

  • Document Intelligence & Search

    Contract analysis, regulatory document processing, research synthesis — structured extraction and semantic Q&A over large private document corpora with source citation.

  • Customer-facing AI Assistant

    Production customer support and sales AI with RAG over product documentation, ticket history and knowledge base — with guardrails, escalation paths and human handoff.

  • Internal Knowledge Base Agent

    Internal AI assistant over company wikis, Notion, Confluence, Slack history and SOPs — with access control per document, evaluation framework and usage analytics.

  • AI-powered Underwriting & Scoring

    Automated analysis of financial documents, credit applications and risk factors — structured data extraction, model scoring integration and human-review queue for edge cases.

  • Code Review & Generation

    AI-assisted code review, documentation generation, test writing and codebase Q&A — fine-tuned on internal standards with guardrails against insecure code suggestions.

  • Autonomous Workflow Agent

    Multi-step autonomous agent for research, data analysis, report generation or process automation — with tool integrations, human-in-the-loop approval gates and comprehensive evaluation.

The AI engineering stack.

Technology selection follows data characteristics, latency requirements, compliance posture and cost model — not hype cycles.

  • Foundation Models

    Stack: OpenAI GPT-4o / GPT-4o-mini, Anthropic Claude 3.5 Sonnet, Mistral Large, Llama 3.1 (Meta), Gemini Pro, AWS Bedrock

  • Orchestration Frameworks

    Stack: LangChain, LlamaIndex, AutoGen, CrewAI, Custom Orchestration, LangGraph

  • Vector Stores

    Stack: Pinecone, Weaviate, pgvector (PostgreSQL), Qdrant, OpenSearch k-NN, Chroma

  • Evaluation & Observability

    Stack: LangSmith, Datadog LLM Observability, Arize Phoenix, Weights & Biases, Custom Eval Pipelines

  • Infrastructure

    Stack: Python / FastAPI, AWS Lambda / ECS, Azure OpenAI Service, Docker, Kubernetes, Redis (semantic cache)

  • Fine-tuning & Training

    Stack: OpenAI Fine-tuning API, Hugging Face PEFT / LoRA, Axolotl, Modal, AWS SageMaker, Together AI

AI systems introduce a new class of security risks that traditional security reviews do not cover.

Prompt injection, data leakage through context and hallucination risk are not bugs. They are fundamental properties of LLMs that require architectural mitigations.

  • Prompt Injection

    Malicious user input can override system instructions, extract confidential context or hijack agent tool calls. Input validation, instruction hierarchy enforcement, output filtering and sandboxed tool execution are required at every user-facing boundary. Don't rely on the model to refuse — enforce at the architecture level.

  • Data Leakage Through Context

    RAG systems surface private documents based on semantic similarity. Without per-user access control at the retrieval layer, users can extract documents they are not authorised to see by crafting queries semantically similar to restricted content. Access control must be enforced at the retrieval layer, not just the document storage layer.

  • PII in RAG Data

    Personal data ingested into vector embeddings and retrieval corpora creates GDPR and privacy obligations. PII must be identified, redacted or tagged before embedding, with deletion procedures that remove both the source document and its vector representations when a deletion request is received.

  • Hallucination Risk

    LLMs confidently produce false information. Every AI output that informs a decision requires source citation, grounding verification against retrieved documents and confidence scoring. High-stakes domains (medical, legal, financial) require human review workflows for outputs below confidence thresholds.

  • API Key Management

    LLM provider API keys grant access to expensive inference capacity and your prompt templates. Keys must be stored in secrets managers (not environment variables or source code), rotated regularly, scoped by service and monitored for anomalous usage patterns.

  • AI Supply Chain Risk

    Third-party model APIs, embedding models and fine-tuning datasets are supply chain dependencies. Model provider behaviour changes, capability regressions and training data contamination affect your system without warning. Evaluation pipelines that catch quality degradation across model updates are the primary defence.

From use case to production AI system.

  1. 01. AI Discovery & Use Case Definition

    Problem framing, success metrics definition, data availability audit, compliance requirements and build-vs-buy decision for model and infrastructure components.

  2. 02. Data Audit & Pipeline Design

    Data source inventory, quality assessment, chunking strategy design, embedding model selection and data pipeline architecture for ingestion and refresh.

  3. 03. Model Selection & RAG Architecture

    Model evaluation against your specific use case, RAG pipeline design, vector store selection, retrieval strategy and context assembly specification.

  4. 04. Agent & Integration Engineering

    Agent architecture (if applicable), tool integration, multi-agent orchestration, memory system design, guardrails implementation and API integration.

  5. 05. Evaluation Framework

    Automated evaluation pipeline — accuracy metrics, hallucination rate, task completion, latency and cost tracking with regression suites and LLM-as-judge where applicable.

  6. 06. Production Deployment

    Production infrastructure deployment, semantic caching, model routing configuration, monitoring dashboards, alerting and load testing at projected query volume.

  7. 07. Monitoring & Optimisation

    Ongoing evaluation monitoring, cost optimisation, model update testing, retrieval quality tuning and feature velocity as the system and underlying models evolve.

Frequently Asked Questions

When should we use RAG vs fine-tuning?

Use RAG when your knowledge base changes frequently, when you need source citations, or when you cannot afford to retrain for every knowledge update. Use fine-tuning when you need to change the model's output format, style or tone rather than its knowledge — for example, teaching it to respond in your brand voice or produce a specific JSON structure. In most production systems, RAG handles dynamic knowledge and fine-tuning handles style/format. The combination is more powerful than either alone.

How do we evaluate whether the AI is actually working?

You need an evaluation pipeline that runs continuously, not a manual spot-check before each release. Define success metrics before building — task completion rate, factual accuracy against ground truth, hallucination rate, latency p50/p95. Run automated evaluations using a test set of representative queries with expected outputs. Use LLM-as-judge for open-ended quality assessment at scale. Track these metrics over time, especially after model updates or prompt changes. Without this, you cannot detect quality degradation.

How do you prevent prompt injection attacks?

Prompt injection prevention requires multiple layers: input validation and sanitisation before injection into prompts, clear separation between system instructions and user input using structured message formats, output filtering that detects instruction-following deviations, sandboxed tool execution that limits what agents can do even if hijacked, and monitoring for anomalous output patterns. Do not rely solely on the model refusing injections — LLMs are not reliable security boundaries.

What is the difference between an AI agent and a chatbot?

A chatbot responds to single-turn queries within a conversation. An AI agent executes multi-step tasks, uses tools (APIs, databases, browsers, code execution), maintains state across steps, makes decisions about which steps to take next and can complete tasks that take minutes to hours. Agents require significantly more engineering — task decomposition logic, tool definitions, error handling, retry strategies and human-in-the-loop gates for high-stakes actions. Most production 'agents' are actually sophisticated chatbots with tool calling.

How do we control AI infrastructure costs?

Cost control is an architectural design problem. Key levers: model tier routing (use GPT-4o-mini for simple queries, GPT-4o for complex), semantic caching (return cached responses for semantically similar queries), context window management (don't send unnecessary context), batch inference for non-real-time tasks, and retrieval precision tuning (retrieve less but better context). Well-designed cost controls typically reduce AI inference cost by 40-70% vs naive implementations while maintaining or improving quality.

Can we run AI models on-premise to keep data private?

Yes — models like Llama 3.1, Mistral Large, Mixtral 8x7B and Phi-3 can be self-hosted on GPU infrastructure. AWS Bedrock and Azure OpenAI offer managed APIs where data stays within your cloud environment without leaving to third-party model providers. Self-hosting requires GPU instance management, model serving infrastructure (vLLM, TGI, Ollama) and ongoing model update management. The performance and capability gap between self-hosted open models and frontier models (GPT-4o, Claude) is closing but not zero.

How do you handle PII in AI systems?

PII handling in AI systems requires: identification and redaction of PII before ingestion into RAG corpora or fine-tuning datasets, per-user access control at the retrieval layer to prevent cross-user data leakage, data retention policies that account for vector embeddings (deleting a source document does not automatically delete its embeddings), and output monitoring for PII appearing in model responses. GDPR right-to-deletion creates an obligation to remove not just source data but all derived vector representations.

What latency can we expect from a production AI system?

Time-to-first-token (TTFT) for streaming responses from OpenAI GPT-4o is typically 300-800ms. Full response latency for a RAG query (retrieval + generation) ranges from 1.5-8 seconds depending on retrieval complexity, context length and model. Semantic caching reduces latency to <50ms for cached queries. If your product requires sub-1-second full responses, this constrains model selection, context length and retrieval strategy significantly. We design latency budgets at the architecture stage.

How do you prevent AI hallucinations?

Hallucinations cannot be eliminated — they are a fundamental property of LLMs. They can be reduced through: retrieval augmentation that grounds responses in source documents, requiring the model to cite specific passages, output validation against structured schemas, confidence scoring and uncertainty quantification, restricting the model's response scope to retrieved context, and evaluation pipelines that measure hallucination rate continuously. For high-stakes decisions, human review queues for low-confidence outputs are the last line of defence.

How long does a production AI system take to build?

A well-scoped RAG system over a private document corpus — ingestion pipeline, retrieval, LLM integration, evaluation framework, frontend interface — takes 8-14 weeks from architecture to production. A multi-tool AI agent with custom integrations takes 12-20 weeks. Fine-tuned model deployment adds 4-6 weeks for dataset preparation, training and evaluation. The evaluation framework and monitoring infrastructure typically take as long as the core system — teams that skip this ship AI systems they cannot measure or maintain.