PROPELOO

GENERATIVE AI / LLM ENGINEERING

Build AI systems that generate, reason and act — not just respond.

PROPELOO engineers generative AI applications — from LLM-powered content pipelines and document generation through multi-agent systems, code generation tools and the evaluation infrastructure that separates AI that works reliably from AI that works most of the time. Generative AI in production is an engineering problem, not a prompt engineering problem.

A generative AI system that is right 90% of the time will be wrong 10% of the time — and in production, that 10% is always at the worst moment.

The gap between a generative AI demo and a production generative AI system is an evaluation problem. How do you know when the output is correct? How do you know when it has drifted? How do you catch hallucinations before they reach users? How do you measure quality across thousands of daily outputs? These are engineering questions that require evaluation frameworks, output validation pipelines, human review queues and monitoring infrastructure. PROPELOO builds generative AI systems with production-grade evaluation built in from the start — because a system you cannot measure is a system you cannot trust.

The full generative AI engineering stack.

Generative AI in production requires a generation pipeline, an evaluation framework, a human review layer and a monitoring system.

System Layers

  • Model & Prompt Layer: LLM selection, prompt engineering, system message design, temperature/parameter tuning
  • Context & Retrieval Layer: RAG pipeline, document processing, vector search, context window management
  • Output Processing Layer: Structured output parsing, format validation, post-processing, output routing
  • Evaluation Layer: Automated quality scoring, human review queue, regression testing, A/B comparison
  • Infrastructure Layer: Streaming, caching, rate limiting, cost control, model fallback

Core Technical Capabilities

  • Content Generation Pipelines

    Automated content generation at scale — blog posts, product descriptions, reports, emails, code documentation. Template-controlled generation with brand voice enforcement, factual grounding and human review for high-stakes outputs.

  • AI Agents & Orchestration

    Multi-step AI agents that plan, use tools (web search, code execution, database queries, API calls) and iterate toward a goal. ReAct and plan-and-execute patterns. Agent memory and context management across multi-turn tasks.

  • Structured Output Generation

    LLM-powered data extraction and transformation — extract structured JSON from unstructured text, classify documents, generate SQL from natural language, transform data between formats. Pydantic validation of LLM output.

  • Code Generation & Review

    Automated code generation for boilerplate, test writing, documentation and refactoring. Code review agents that check for security issues, style violations and logical errors. Integration with existing CI/CD pipelines.

  • Multimodal Applications

    Applications combining text, image and document understanding — document OCR and extraction with GPT-4 Vision, image description for accessibility, video transcript processing and mixed-media content generation.

  • Evaluation & Quality Infrastructure

    RAGAS for RAG system evaluation, LLM-as-judge for open-ended output quality scoring, regression testing suites for prompt changes, human evaluation workflows and production quality monitoring dashboards.

How we think about generative AI engineering.

Generative AI is non-deterministic. You cannot unit test it. You can only evaluate it statistically — and that requires designing evaluation infrastructure before deploying to production.

  • Evaluation before shipping

    Defining what good output looks like, building a test set of representative inputs with expected outputs, and measuring quality before launch is not optional for production generative AI. Without evaluation: you do not know if a prompt change made things better or worse, you cannot detect model drift when the LLM provider updates, and you cannot measure the impact of your improvements. LangSmith, Arize Phoenix and RAGAS provide the evaluation infrastructure.

    Axiom:

  • Structured output is more reliable than unstructured

    Asking an LLM to return JSON with specific fields is more reliable than parsing JSON from prose output. OpenAI function calling and JSON mode, Anthropic tool use and Pydantic validation schemas constrain the output format and make parsing deterministic. Downstream systems that depend on LLM output should receive validated, typed data — not raw text to parse.

    Axiom:

  • Human review is part of the architecture

    Not every generated output should reach users without review. High-stakes outputs (legal documents, medical information, financial advice) should have human review queues. Lower-stakes outputs can have sampling-based review. Build the human review workflow into the system architecture — not as an afterthought for when things go wrong.

    Axiom:

  • Cost control is a system requirement

    GPT-4o at $5/1M tokens is cheap per query and expensive at scale. 1M daily queries = $5,000/day in LLM costs alone. Cost optimisation: use smaller, cheaper models (GPT-4o-mini, Claude Haiku) for tasks that do not require full capability, cache deterministic outputs, batch non-urgent requests, and set per-user rate limits. Cost monitoring with alerts is required infrastructure for production generative AI.

    Axiom:

The generative AI system decisions that matter.

Each choice affects quality, cost, latency and reliability.

  • Which LLM for which task?

    Impact: Match model to task. GPT-4o-mini for classification, extraction and simple generation. GPT-4o / Claude 3.5 for complex reasoning, code generation and nuanced content. Self-hosted for privacy requirements or very high volume.

    • GPT-4o — best instruction following, highest cost, best default for complex tasks
    • Claude 3.5 Sonnet — strong reasoning, 200K context, best for long documents
    • GPT-4o-mini — 10x cheaper, 80% of quality, right for high-volume simple tasks
    • Llama 3 / Mistral (self-hosted) — no API cost, full privacy, requires GPU infrastructure
  • Streaming vs batch?

    Impact: Streaming improves perceived UX for interactive applications. Batch is simpler for document generation workflows. Choose based on whether the user is waiting for the response in real time.

    • Streaming (SSE) — perceived latency improvement, user sees output as it generates
    • Batch request — simpler, full response before displaying, higher apparent latency
    • Background job — async generation, user notified when complete
    • Hybrid by use case — streaming for chat, batch for document generation
  • Caching strategy?

    Impact: Semantic caching for FAQ-style applications can reduce LLM API costs by 30-60%. For creative or personalised generation, caching is inappropriate. Exact match caching is valuable for deterministic queries like classification.

    • No caching — simplest, every request hits LLM API
    • Exact match cache — same input = cached response, low hit rate
    • Semantic cache (GPTCache) — similar inputs return cached response, higher hit rate
    • Partial caching (cache prefix) — cache common system prompt prefixes
  • Output validation approach?

    Impact: Pydantic + function calling / JSON mode for any application that depends on structured LLM output. LLM-as-judge for quality validation of open-ended text generation. Never rely on LLM output format without validation in production.

    • No validation — fastest, unreliable output format
    • Pydantic schema + retry — parse and validate, retry on failure
    • OpenAI JSON mode / function calling — constrained output format, reliable parsing
    • LLM-as-judge validation — use second LLM to validate output quality
  • Multi-agent vs single-agent?

    Impact: Start with the simplest approach that solves the problem. Multi-agent systems are powerful but add orchestration complexity, cost and failure modes. ReAct agents for tasks requiring tool use and iteration. Pipelines for well-defined multi-step workflows.

    • Single LLM call — simplest, lowest cost, limited to single-step reasoning
    • Chain of thought (single model, multi-step prompt) — better reasoning, same model
    • Multi-agent pipeline — specialist agents per subtask, orchestrated workflow
    • ReAct agent (reason + act loop) — tool use, iterative refinement, highest capability
  • Evaluation framework?

    Impact: LLM-as-judge scoring for open-ended generation quality is the most scalable approach with acceptable accuracy. RAGAS for RAG systems. Always validate automated evaluation against human judgment on a calibration set before trusting it for regression detection.

    • Manual review only — slow, does not scale, no regression detection
    • Automated metrics (BLEU/ROUGE) — fast, poor correlation with actual quality for LLM outputs
    • LLM-as-judge — flexible, good correlation with human judgment, adds cost
    • RAGAS (for RAG) — specialized RAG evaluation, covers retrieval + generation

What PROPELOO builds.

  • Content Generation Platform

    Automated blog post, product description and marketing copy generation with brand voice enforcement, factual grounding and human review workflow for high-stakes content.

  • Document Processing Pipeline

    Extract structured data from unstructured documents — contracts, invoices, medical records, financial statements — with Pydantic validation and confidence scoring.

  • AI Coding Assistant

    Code generation, review and documentation tool integrated into developer workflow — PR review automation, test generation, documentation writing and boilerplate code generation.

  • Multi-agent Research System

    Autonomous research agent that searches the web, reads documents, synthesises findings and produces structured reports with citations — with human review for accuracy-critical outputs.

  • AI-powered Customer Emails

    Personalised customer communication generation at scale — support responses, follow-up emails, offer recommendations — with brand voice, tone matching and human review sampling.

  • Data Transformation Pipeline

    LLM-powered ETL pipeline that understands messy, unstructured data sources and transforms them into clean, structured database records with validation and confidence scoring.

The generative AI stack.

Model APIs, orchestration, evaluation and infrastructure each require specific tooling.

  • LLM APIs

    Stack: OpenAI GPT-4o/mini, Anthropic Claude 3.5, Google Gemini 1.5, Groq (fast inference), Together AI (open source)

  • Orchestration

    Stack: LangChain, LlamaIndex, LangGraph (agents), Haystack, Custom Python pipeline

  • Output Validation

    Stack: Pydantic v2, Instructor (structured output), Guardrails AI, OpenAI function calling

  • Evaluation

    Stack: LangSmith, RAGAS, Arize Phoenix, PromptFoo, Custom eval suite

  • Caching & Cost

    Stack: GPTCache, Redis (semantic cache), LangSmith (cost tracking), Helicone

  • Infrastructure

    Stack: FastAPI, Celery (async jobs), AWS Lambda, PostgreSQL, Redis

Generative AI security in production.

LLM applications introduce security risks that traditional application security does not cover.

  • Prompt Injection

    Users can craft inputs that override system instructions. For agentic systems with tool access, prompt injection can trigger unauthorised tool calls, data exfiltration or privilege escalation.

  • Data Exfiltration via LLM

    Multi-tenant systems where user inputs are included in prompts alongside other users' data can leak information across tenants. Strict context isolation and per-user data sandboxing in the prompt context are required.

  • PII in LLM Calls

    User inputs containing PII sent to third-party LLM APIs may violate GDPR, HIPAA or CCPA. PII detection and redaction before LLM API calls, data processing agreements with LLM providers and self-hosted models for regulated data are mitigation options.

  • API Key Management

    LLM API keys with no spend limits or rate limiting will be exploited if leaked. Server-side key management, per-user rate limiting and spend alerts are required for production systems.

  • Output Sanitisation

    Generated content must be sanitised before rendering in web interfaces — LLM-generated HTML or JavaScript can contain XSS vectors if rendered unsafely. Treat all LLM output as untrusted user input.

  • Model Dependency Risk

    Single-provider LLM dependency means provider outages take down your product. Multi-provider fallback (OpenAI → Anthropic → self-hosted) with automatic failover is required for production systems with uptime SLAs.

From prototype to production AI.

  1. 01. Use Case Definition

    Define specific generation task, quality criteria, acceptable failure modes and evaluation metrics before any implementation.

  2. 02. Model & Architecture Selection

    LLM selection, prompt strategy, output format, caching approach and evaluation framework design.

  3. 03. Pipeline Development

    Generation pipeline, structured output validation, error handling and retry logic.

  4. 04. Evaluation Framework

    Test set construction, LLM-as-judge setup, quality baselines and regression test suite.

  5. 05. Human Review Workflow

    Review queue for high-stakes outputs, sampling strategy for quality monitoring, feedback collection.

  6. 06. Production Integration

    API development, streaming, rate limiting, cost controls and monitoring dashboard.

  7. 07. Monitoring & Iteration

    Quality drift detection, cost monitoring, model performance comparison and iterative improvement.

Frequently Asked Questions

How do we measure quality of generated content?

Traditional metrics (BLEU, ROUGE) correlate poorly with human judgment for LLM outputs. LLM-as-judge — using GPT-4 or Claude to score outputs against defined criteria — correlates well with human evaluation and scales to thousands of outputs per day. RAGAS provides specialised metrics for RAG systems. Always validate your evaluation methodology against human judgments on a calibration set before using it for production monitoring.

How do we reduce hallucination?

RAG grounding is the primary mechanism — constrain generation to retrieved, verified content. Structured output with source citation. Temperature close to 0 for factual tasks. "I don't know" prompting for low-confidence responses. LLM-as-judge hallucination detection on outputs before delivery. No mechanism eliminates hallucination entirely — monitoring and human review for high-stakes outputs remain necessary.

What is an AI agent and when do we need one?

An AI agent is an LLM that can use tools (search, code execution, database queries, API calls) and iterate over multiple steps to complete a task. Use agents when: the task requires information that is not in the context window (needs search), requires computation (needs code interpreter), or requires multiple sub-tasks that depend on each other. Do not use agents when a single well-designed prompt or RAG call can solve the problem — agents add complexity, latency and failure modes.

How much does production generative AI cost?

At 10,000 daily queries with average 2,000 tokens per query: GPT-4o costs ~$100/day, Claude 3.5 Sonnet ~$90/day, GPT-4o-mini ~$10/day. Self-hosted Llama 3 70B on AWS: ~$50/day for a single A100 instance. For most applications, GPT-4o-mini or Claude Haiku is 80% of GPT-4o quality at 10% of the cost. The right architecture uses tiered models — complex tasks get the expensive model, simple classification and extraction get the cheap model.