PROPELOO

KUBERNETES / CONTAINER ORCHESTRATION

Run containers at production scale without the operational surprises.

PROPELOO engineers Kubernetes infrastructure — from EKS/GKE cluster design and workload migration through Helm chart development, autoscaling, RBAC, network policies and the observability stack that makes a K8s cluster operable by your team at 2am. Kubernetes is not a deployment target. It is a platform with its own engineering requirements.

A Kubernetes cluster that nobody on your team understands is not infrastructure — it is a liability waiting to become an incident.

Kubernetes reduces operational complexity for large-scale container workloads. It does not eliminate complexity — it moves it. The complexity of managing individual servers becomes the complexity of managing cluster configuration, RBAC, network policies, resource requests/limits, pod disruption budgets, horizontal pod autoscaling, ingress controllers and the etcd backup schedule. Every one of these has a failure mode. PROPELOO builds Kubernetes infrastructure that is documented, version-controlled in Terraform and Helm, monitored end-to-end with Prometheus/Grafana, and operated according to runbooks that any engineer on your team can follow — not just the one who set it up.

The full Kubernetes engineering stack.

Production Kubernetes requires engineering across six domains before the first workload is deployed.

System Layers

  • Cluster Infrastructure Layer: EKS/GKE/AKS provisioning, node groups, networking (VPC CNI), control plane configuration
  • Workload Layer: Deployment manifests, StatefulSets, DaemonSets, Jobs/CronJobs, resource requests/limits
  • Networking Layer: Ingress controller, service mesh, network policies, DNS, load balancer configuration
  • Security Layer: RBAC, Pod Security Standards, network policies, secrets management, image scanning
  • Observability Layer: Metrics (Prometheus), logs (Loki), traces (Tempo), alerts (Alertmanager), dashboards (Grafana)

Core Technical Capabilities

  • Cluster Design & Provisioning

    EKS/GKE/AKS cluster in Terraform — node group sizing, auto-scaling groups, spot/on-demand mix, networking (VPC CNI vs Calico), cluster version strategy and multi-AZ deployment for HA.

  • Workload Engineering

    Deployment manifests with correct resource requests/limits, liveness/readiness/startup probes, PodDisruptionBudgets, affinity/anti-affinity rules, topologySpreadConstraints and HorizontalPodAutoscaler configuration.

  • Helm Chart Development

    Production Helm charts with configurable values, environment-specific overrides, secrets management via External Secrets Operator or Vault Agent, and chart testing with helm-unittest.

  • GitOps with ArgoCD

    ArgoCD ApplicationSets for multi-environment GitOps, sync policies, health checks, rollback triggers and progressive delivery (Argo Rollouts) for blue/green and canary deployments.

  • Cluster Security Hardening

    RBAC with least-privilege service accounts, Pod Security Standards (restricted profile), NetworkPolicies for east-west traffic segmentation, OPA/Gatekeeper admission policies and Falco for runtime threat detection.

  • Autoscaling Architecture

    HorizontalPodAutoscaler (CPU/memory/custom metrics), VerticalPodAutoscaler for right-sizing, Cluster Autoscaler for node-level scaling, KEDA for event-driven autoscaling and Karpenter for cost-optimised node provisioning.

How we think about Kubernetes.

Kubernetes is justified by specific requirements: independent scaling of different services, sophisticated deployment strategies, or team autonomy at scale. Without those requirements, ECS Fargate or Cloud Run provides 80% of the capability with 20% of the operational overhead.

  • Resource requests and limits are not optional

    A pod without resource requests gets scheduled on any node regardless of available resources — it will evict other pods when the node runs out of memory. A pod without resource limits can consume unlimited CPU, starving neighbouring pods. Every production workload must have correctly tuned resource requests (based on actual usage data) and limits (set above expected peak, below what would impact neighbours). This is the most common cause of mysterious production instability in K8s.

    Axiom:

  • Cluster autoscaling must be matched to workload patterns

    Cluster Autoscaler adds nodes when pods are unschedulable and removes them when they are underutilised. If your workload spikes predictably (batch jobs at midnight, traffic peaks during business hours), node scale-up latency (2-3 minutes) may be too slow. Karpenter provides faster scale-up with more flexible node provisioning. Pre-scaling nodes before expected traffic spikes using scheduled HPA adjustments is the practical solution for predictable load patterns.

    Axiom:

  • RBAC must be designed before workloads are deployed

    Retrofitting RBAC onto a running cluster requires auditing every service account and every role binding that was created with convenience in mind. Design RBAC from day one: each application service account has only the permissions it needs, cluster-admin is granted to no application workloads, and human access is via short-lived tokens rather than permanent kubeconfig credentials.

    Axiom:

  • Runbooks make clusters operable

    A Kubernetes cluster with excellent Prometheus dashboards but no runbooks creates alert fatigue — engineers see the alert, do not know what to do, and start ignoring them. Every alert must have a corresponding runbook that answers: what does this alert mean, what are the likely causes, what are the investigation steps, and what is the remediation. Runbooks stored in the same git repository as the manifests they document.

    Axiom:

The Kubernetes decisions that define operational complexity.

Each choice has compounding operational implications.

  • EKS vs GKE vs AKS vs self-hosted?

    Impact: Use managed K8s (EKS/GKE/AKS) in virtually all cases. Self-hosted clusters require your team to manage etcd backups, control plane upgrades and cluster certificates — operational overhead that managed K8s eliminates.

    • EKS (AWS) — managed control plane, deep AWS integration, best if primary cloud is AWS
    • GKE (GCP) — most mature managed K8s, Autopilot mode reduces node management
    • AKS (Azure) — managed control plane, deep Azure AD integration, best for Microsoft-heavy orgs
    • Self-hosted (kubeadm) — full control, full operational burden, only for specific compliance requirements
  • Helm vs Kustomize vs plain manifests?

    Impact: Helm for third-party applications (Prometheus, cert-manager, ingress-nginx) — the Helm chart ecosystem is unmatched. Kustomize for your own application manifests — simpler than Helm for apps you control. Plain manifests only for single-environment setups.

    • Helm — templating, package management, large chart ecosystem, versioned releases
    • Kustomize — overlay-based, no templating, built into kubectl, good for environment variants
    • Plain manifests — simplest, no abstraction, difficult to manage across environments
    • Helm + Kustomize — use Helm for third-party charts, Kustomize for your own apps
  • ArgoCD vs Flux for GitOps?

    Impact: ArgoCD for most teams — the UI provides excellent visibility into sync status, drift and deployment history. Flux for teams that prefer a more Kubernetes-native, operator-based approach.

    • ArgoCD — excellent UI, application-centric model, good for multi-tenant clusters
    • Flux — more Kubernetes-native, lighter weight, better for simple setups
    • Manual kubectl apply — no GitOps, drift detection impossible, not recommended for production
    • Spinnaker — complex, more features than most teams need
  • Service mesh (Istio vs Linkerd vs none)?

    Impact: No service mesh for clusters under 10 services. Linkerd for teams that need mTLS and basic traffic management without Istio complexity. Istio for advanced traffic management (circuit breaking, fault injection, weighted routing) requirements.

    • No service mesh — services handle their own resilience, simpler to operate
    • Istio — comprehensive, mTLS, traffic management, observability, significant resource overhead
    • Linkerd — lighter than Istio, simpler, mTLS by default, less features
    • Cilium — eBPF-based, handles both CNI and service mesh, growing adoption
  • Node sizing strategy?

    Impact: Karpenter for AWS EKS is increasingly the standard — it provisions nodes of exactly the right size for pending workloads, significantly reducing cluster costs. For GKE, GKE Autopilot provides similar optimisation.

    • Few large nodes — simpler to manage, higher blast radius per node failure
    • Many small nodes — better isolation, higher bin-packing efficiency, more failure domains
    • Mixed (large for stateful, small for stateless) — optimised for workload types
    • Karpenter with auto-sizing — nodes provisioned per workload, optimal cost, AWS/GCP only
  • Secrets management?

    Impact: External Secrets Operator + AWS Secrets Manager is the production standard — secrets live in a dedicated secret store with rotation capability, audit trail and fine-grained access control. Never store base64-encoded Kubernetes Secrets in git.

    • Kubernetes Secrets (base64) — simplest, unencrypted at rest by default, version-controlled risk
    • Sealed Secrets — encrypt secrets for git storage, decrypt in cluster
    • External Secrets Operator + AWS Secrets Manager/Vault — secrets from external store, rotatable
    • Vault Agent sidecar — Vault-native, injected as sidecar, complex setup

What PROPELOO builds.

  • EKS Cluster from Scratch

    Complete EKS setup in Terraform — VPC, node groups, IRSA, cluster autoscaler, ingress-nginx, cert-manager, External Secrets, Prometheus stack and ArgoCD.

  • Workload Migration to K8s

    Migrate from EC2, ECS or Heroku to Kubernetes — containerise applications, write Helm charts, set resource requests/limits, configure HPA and validate with load testing.

  • K8s Security Hardening

    Security audit of existing cluster — RBAC review, Pod Security Standards enforcement, NetworkPolicy implementation, OPA admission policies and Falco runtime monitoring.

  • GitOps Implementation

    ArgoCD setup with ApplicationSets for multi-environment GitOps — staging and production promotion workflow, sync policies, health checks and rollback automation.

  • Autoscaling Architecture

    HPA with custom metrics (KEDA), Karpenter for node autoscaling, spot instance configuration and load testing to validate scale-up behaviour.

  • Multi-cluster Architecture

    Multi-region K8s with ArgoCD ApplicationSets for cross-cluster deployment, cluster federation for shared services and region failover strategy.

The Kubernetes engineering stack.

Cluster provisioning, workload management and observability each require specific tooling.

  • Cluster Provisioning

    Stack: Terraform (EKS/GKE/AKS), eksctl, Pulumi, Cluster API

  • Workload Management

    Stack: Helm 3, Kustomize, ArgoCD, Flux v2, Argo Rollouts

  • Networking

    Stack: ingress-nginx, AWS Load Balancer Controller, Istio, Linkerd, Cilium, cert-manager

  • Security

    Stack: OPA/Gatekeeper, Falco, Trivy Operator, External Secrets Operator, Sealed Secrets

  • Autoscaling

    Stack: HPA, VPA, Cluster Autoscaler, Karpenter, KEDA

  • Observability

    Stack: kube-prometheus-stack, Grafana, Loki, Tempo, OpenTelemetry Operator

Kubernetes security requires hardening at every layer.

Default K8s configurations prioritise ease of use over security. Production clusters require explicit hardening.

  • RBAC Least Privilege

    Default service accounts should have no permissions. Every application gets a dedicated service account with only the cluster permissions it needs. ClusterAdmin should never be granted to application workloads. Human access via short-lived tokens (AWS SSO / GCP Workload Identity) rather than permanent kubeconfig.

  • Pod Security Standards

    Kubernetes Pod Security Standards (restricted profile) prevent containers running as root, privilege escalation, host network access and writable root filesystems. Applied via namespace labels. Pods that cannot meet restricted profile require explicit justification and baseline profile exceptions.

  • Network Policies

    By default, all pods in a cluster can communicate with all other pods. NetworkPolicies define explicit allow rules — a service should only accept traffic from pods that legitimately need to call it. Default-deny ingress and egress policies applied to all namespaces, with explicit allow rules per application.

  • Secrets Encryption

    Kubernetes Secrets are base64-encoded (not encrypted) at rest by default. Enable encryption at rest for etcd using a KMS provider. Use External Secrets Operator to pull secrets from AWS Secrets Manager or HashiCorp Vault at runtime rather than storing secrets in the cluster.

  • Container Image Security

    Trivy Operator scans running containers for CVEs and reports violations via Kubernetes resources. Admission policies (OPA/Gatekeeper) block deployment of images with critical CVEs, images without digest pinning and images from untrusted registries.

  • Control Plane Access

    Kubernetes API server access must be restricted — private endpoint only (not publicly accessible), access via VPN or bastion, short-lived credentials and audit logging for all API server requests. CloudTrail/Cloud Audit Logs for all EKS/GKE API calls.

From zero to production K8s cluster.

  1. 01. Architecture Design

    Cluster topology, node group strategy, networking, secrets management, GitOps approach and observability stack selection.

  2. 02. Cluster Provisioning

    EKS/GKE/AKS in Terraform — VPC, node groups, IAM/IRSA, cluster autoscaler and base add-ons.

  3. 03. Platform Add-ons

    ingress-nginx, cert-manager, External Secrets Operator, Prometheus stack, Loki and ArgoCD deployed via Helm/ArgoCD.

  4. 04. Security Hardening

    RBAC design, Pod Security Standards, NetworkPolicies, OPA admission policies and Falco runtime monitoring.

  5. 05. Workload Migration

    Application containerisation, Helm chart development, resource tuning, HPA configuration and staged rollout.

  6. 06. Observability

    Prometheus dashboards, Loki log aggregation, distributed tracing, alerting rules and runbook documentation per alert.

  7. 07. Runbooks & Handoff

    Cluster operations runbook, common failure scenarios, upgrade procedures and team knowledge transfer.

Frequently Asked Questions

When does Kubernetes make sense vs ECS Fargate?

Kubernetes for: systems with many services (10+) where Kubernetes scheduling intelligence, advanced deployment strategies (canary, blue/green), and ecosystem tooling (service mesh, KEDA, Argo Rollouts) provide real value. ECS Fargate for: systems with fewer services, teams without existing Kubernetes expertise, or AWS-native organisations where ECS integration with IAM, CloudWatch and ALB is a productivity advantage. ECS is not a stepping stone to Kubernetes — it is a legitimate production choice for the right use case.

How do we set resource requests and limits correctly?

Measure actual resource usage first — VPA Recommender in recommendation mode observes pod resource usage and suggests requests/limits based on historical data. Set requests at the P50 (median) CPU and P95 memory usage. Set limits at 2-4x requests for CPU (burstable), and at P99 memory usage plus 20% buffer for memory (OOM kill on limit breach is instantaneous and painful). Never set memory limits lower than memory requests. Test limits by load testing to the expected peak before applying to production.

What is KEDA and when should we use it?

KEDA (Kubernetes Event-driven Autoscaling) scales pods based on event sources — queue depth (SQS, Kafka), HTTP request rate, database query count, custom metrics. Standard HPA scales on CPU and memory. KEDA scales on what actually drives your load. Use KEDA for: batch processing services that should scale from 0 when the queue is empty, workers that process Kafka messages and should scale with lag, and services where CPU/memory are poor proxies for actual load.

What is Karpenter and how is it different from Cluster Autoscaler?

Cluster Autoscaler scales node groups — it adds/removes nodes from pre-defined node groups. Karpenter provisions individual nodes of exactly the right size for pending workloads — it can provision a c6g.8xlarge for a workload that needs it without maintaining a pre-defined node group of that instance type. Karpenter is faster (sub-60-second node provisioning vs 2-3 minutes for CA), more cost-efficient (right-size per workload) and supports more flexible spot/on-demand mix strategies. Available for EKS and GKE.

How do we do zero-downtime deployments in Kubernetes?

Prerequisites: correct readiness probes (pod only receives traffic when actually ready), PodDisruptionBudget (minimum available pods during disruption), resource requests (scheduler places pods on nodes with capacity). Then: rolling update strategy with maxUnavailable: 0 and maxSurge: 1 ensures old pods are not removed until new pods are ready. For more control: Argo Rollouts for blue/green (instant cutover and rollback) or canary (gradual traffic shift). Pre-stop hook with sleep to drain in-flight requests before pod termination.