PROPELOO

AI CLOUD INFRASTRUCTURE / ML PLATFORM

Build the infrastructure that makes AI models fast, cheap and reliable in production.

PROPELOO engineers AI cloud infrastructure — from model serving and GPU cluster management through inference optimisation, model registry, feature stores and the MLOps pipeline that automates training, evaluation and deployment. Running an AI model in a Jupyter notebook is not production. Production AI infrastructure is a distinct engineering discipline.

A model that takes 30 seconds to respond in production was never production-ready — it was demo-ready.

ML model serving has unique infrastructure requirements: GPU memory management, batching to amortise per-request overhead, model warmup to eliminate cold starts, auto-scaling that responds to inference load patterns, and continuous deployment pipelines that can roll back a model version without deploying new application code. These requirements are not met by simply deploying a model in a Docker container. PROPELOO builds AI infrastructure where inference is fast (optimised serving, batching, caching), cheap (right-sized GPU instances, spot instance strategies, model quantisation), reliable (health checks, autoscaling, fallback endpoints) and observable (request latency, GPU utilisation, model performance metrics).

The AI cloud infrastructure stack.

System Layers

  • Serving Layer: Model server (vLLM, TorchServe, Triton), batching, GPU memory management, endpoint routing
  • Scaling Layer: GPU autoscaling, spot instances, queue-based scaling (KEDA), cold start management
  • MLOps Layer: Training pipelines, model registry, experiment tracking, automated evaluation, deployment gates
  • Data Layer: Feature store, training data pipelines, data versioning, offline/online feature serving
  • Observability Layer: Model latency, GPU utilisation, model drift detection, prediction quality monitoring

Core Technical Capabilities

  • Model Serving

    vLLM for LLM serving (10-30x throughput vs naive serving via continuous batching), TorchServe for PyTorch models, NVIDIA Triton for multi-framework serving, FastAPI custom inference server.

  • Inference Optimisation

    Model quantisation (FP16, INT8, INT4 with GPTQ/AWQ), TensorRT compilation for NVIDIA GPUs, ONNX export for cross-platform deployment, KV cache optimisation for LLMs.

  • MLOps Pipeline

    MLflow or Weights & Biases for experiment tracking, model registry with versioning, automated evaluation pipelines (accuracy, latency, bias), CI/CD for model deployment with rollback.

  • GPU Infrastructure

    EKS with GPU node groups (g4dn, p3, p4), GPU operator for device plugin, KEDA for queue-based autoscaling, spot instances for cost optimisation, Karpenter for fast GPU node provisioning.

  • Feature Store

    Feast or Tecton feature store — offline feature engineering pipeline, online feature serving for low-latency inference, feature versioning and point-in-time correctness for training.

  • Model Monitoring

    Prediction latency (p50/p95/p99), GPU utilisation and memory, model drift detection (distribution shift), data quality monitoring, business metric correlation.

How we think about AI infrastructure.

AI infrastructure is not harder than regular infrastructure — it is differently hard. The unique challenges are GPU management, batching for efficiency, and the model versioning complexity that does not exist for stateless services.

  • vLLM changes LLM serving economics

    Naive LLM serving (one request at a time) leaves GPU memory largely idle between requests and does not amortise the KV cache overhead across requests. vLLM's continuous batching fills GPU memory with multiple concurrent requests, achieving 10-30x throughput improvement. Any LLM serving infrastructure that is not using continuous batching is leaving GPU capacity on the table.

    Axiom:

  • Quantisation trades quality for cost

    FP16 inference is ~50% memory savings vs FP32. INT8 is ~75%. INT4 (GPTQ/AWQ) is ~87.5%. A 70B parameter model in FP16 requires ~140GB GPU memory (4x A100 80GB). In INT4: ~35GB (single A100). Quantisation typically introduces 1-3% accuracy degradation. For most production applications, INT8 is the sweet spot — significant memory savings, negligible quality impact.

    Axiom:

  • Cold starts are UX killers for inference

    A GPU instance that was scaled to zero takes 3-5 minutes to start, warmup and serve the first request. For interactive applications, this is unacceptable. Minimum instance count of 1 (never scale to zero), provisioned capacity for baseline load, and scale-out time hidden behind request queuing are the production patterns.

    Axiom:

  • Model rollback must be independent of application rollback

    Deploying a new model version should not require a new application deployment. The model registry, model server and application code must be independently versioned and deployable. A model version that degrades accuracy needs to be rolled back in minutes, not hours.

    Axiom:

AI infrastructure decisions.

  • Model serving framework?

    Impact: vLLM for LLM serving — the continuous batching throughput improvement justifies the setup complexity. TorchServe for traditional ML models. FastAPI for simple models where optimised serving is not justified.

    • vLLM — best for LLMs, continuous batching, OpenAI-compatible API
    • TorchServe — PyTorch-native, multi-model serving, batching
    • NVIDIA Triton — multi-framework, high-performance, complex setup
    • FastAPI custom — full control, less optimised, fastest to start
  • GPU instance strategy?

    Impact: Reserved instances for baseline GPU capacity, spot instances for scale-out. Critical inference endpoints always have at least one reserved instance. Spot for training and batch inference.

    • On-demand GPU — highest cost, immediate availability
    • Reserved GPU instances — 30-60% savings, 1-year commitment
    • Spot GPU instances — 70-90% savings, interruption risk
    • Hybrid (reserved baseline + spot scale-out) — best cost/reliability balance
  • Autoscaling trigger?

    Impact: KEDA with queue depth for inference autoscaling — scale GPU workers based on pending request queue length. Responds to demand faster than utilisation-based scaling.

    • CPU/GPU utilisation (HPA) — lags actual inference demand
    • Request queue depth (KEDA) — responds to demand directly
    • Custom metric (requests/second) — precise, requires custom metrics
    • No autoscaling — fixed capacity, simple, expensive
  • MLOps platform?

    Impact: MLflow for most teams — open source, self-hosted, covers experiment tracking and model registry without per-user cost. W&B for teams that prioritise collaboration features.

    • MLflow — open source, experiment tracking, model registry, self-hosted
    • Weights & Biases (W&B) — managed, excellent UX, cost at scale
    • Kubeflow — Kubernetes-native, comprehensive, complex
    • AWS SageMaker — managed, AWS-specific, expensive
  • Feature store?

    Impact: Feast for ML teams that need point-in-time correct features for training and low-latency online serving. No feature store acceptable only for models with no real-time features.

    • Feast — open source, offline + online serving
    • Tecton — managed, production-grade, expensive
    • Redis only — simple online features, no offline training consistency
    • No feature store — acceptable for simple models
  • Training infrastructure?

    Impact: Kubeflow pipelines for reproducible, orchestrated training workflows. Single GPU for models that fit in memory. Distributed training only when model or data size genuinely requires it.

    • Single large GPU instance — simple, limited parallelism
    • Distributed training (PyTorch DDP) — multi-GPU, complex setup
    • Kubeflow pipelines — orchestrated multi-step training
    • AWS SageMaker training jobs — managed, expensive at scale

What PROPELOO builds.

  • LLM Inference Platform

    vLLM-based LLM serving on GPU nodes — OpenAI-compatible API, continuous batching, autoscaling and cost-optimised spot instance strategy.

  • ML Model Serving Platform

    TorchServe/Triton multi-model serving with versioning, A/B testing, shadow deployment and automated rollback.

  • MLOps Pipeline

    End-to-end MLOps with MLflow — training pipeline, experiment tracking, model registry, evaluation gates and automated deployment.

  • Fine-tuning Infrastructure

    GPU cluster for LLM fine-tuning — LoRA/QLoRA setup, distributed training, checkpoint management and evaluation pipeline.

  • AI Cost Optimisation

    Audit existing AI infrastructure costs — model quantisation, batching optimisation, spot instance migration and right-sizing. Typically 40-60% cost reduction.

  • Feature Store & Training Pipeline

    Feast feature store with offline training pipeline, online serving for real-time inference and point-in-time correctness.

The AI infrastructure stack.

  • Model Serving

    Stack: vLLM, TorchServe, NVIDIA Triton, FastAPI (custom), Ray Serve

  • Optimisation

    Stack: TensorRT, ONNX Runtime, GPTQ (quantisation), AWQ (quantisation), Flash Attention

  • MLOps

    Stack: MLflow, Weights & Biases, DVC (data versioning), Kubeflow Pipelines, ZenML

  • Infrastructure

    Stack: EKS with GPU nodes, Karpenter, KEDA (queue autoscaling), NVIDIA GPU Operator, Spot instance configuration

  • Features

    Stack: Feast (feature store), Redis (online features), Apache Spark (offline processing), dbt

  • Monitoring

    Stack: NVIDIA DCGM (GPU metrics), Prometheus + Grafana, Arize Phoenix, Evidently AI (drift)

AI infrastructure security for production ML systems.

  • Model access control

    Model endpoints require authentication — API keys or JWT. Internal model servers behind VPC boundary with no public internet access. GPU nodes in private subnets.

  • Training data security

    Training data often contains sensitive information. Encryption at rest in S3, access controls per dataset, data lineage tracking and deletion capability for GDPR compliance.

  • Model IP protection

    Fine-tuned model weights are valuable intellectual property. Encrypted model storage, access logging for model downloads, and model API serving (never expose weights to clients).

  • Inference request logging

    Inference requests may contain sensitive user data. Log at appropriate granularity — metadata yes, PII no. Retention limits. Access controls on inference logs.

  • Supply chain security

    ML models and dependencies have supply chain risks. Verify model checksums from Hugging Face, scan Python packages for vulnerabilities, pin dependency versions.

  • GPU instance security

    GPU instances have the same security requirements as CPU instances plus: prevent model exfiltration via shared GPU memory, secure model loading process, least-privilege IAM for GPU worker service accounts.

From Jupyter to production AI infrastructure.

  1. 01. Infrastructure Design

    Model serving requirements, GPU sizing, autoscaling strategy, MLOps tooling selection.

  2. 02. GPU Cluster Setup

    EKS GPU node groups, GPU operator, KEDA, Karpenter, spot instance configuration.

  3. 03. Model Serving

    vLLM/TorchServe deployment, batching configuration, health checks, OpenAI-compatible API.

  4. 04. Optimisation

    Quantisation, TensorRT compilation, latency benchmarking, cost analysis.

  5. 05. MLOps Pipeline

    MLflow setup, training pipeline, model registry, evaluation gates, deployment automation.

  6. 06. Monitoring

    GPU metrics, inference latency, model quality dashboards, drift detection.

  7. 07. Cost Optimisation

    Spot instance migration, reserved capacity strategy, quantisation ROI analysis.

Frequently Asked Questions

What is vLLM and why is it better than naive serving?

vLLM is an LLM inference server that uses PagedAttention — a technique that manages KV cache memory like an OS page table, allowing multiple requests to share GPU memory efficiently. Combined with continuous batching (dynamically grouping in-flight requests), vLLM achieves 10-30x throughput compared to serving one request at a time. It also exposes an OpenAI-compatible API, making migration from the OpenAI API straightforward.

How do we reduce GPU inference costs?

Four main levers: (1) Quantisation (INT8/INT4) reduces memory requirements 2-4x, enabling smaller/cheaper GPU instances. (2) vLLM continuous batching increases GPU utilisation from ~20% to ~80% for LLMs. (3) Spot instances (70-90% cheaper than on-demand) for stateless inference with request retry on interruption. (4) Right-sizing: many teams run inference on oversized GPU instances. A/100 is not always necessary when A10G handles the model.

What is model drift and how do we detect it?

Model drift occurs when model predictions degrade over time because the input data distribution changes (data drift) or the relationship between inputs and outputs changes (concept drift). Detection: monitor prediction distribution statistics (mean, variance, distribution shape) against training baseline using Evidently AI or Arize Phoenix. Alert when statistical tests indicate significant deviation. For classification models: monitor accuracy against a labelled holdout sample updated monthly.