Frequently Asked Questions
How long does it take to set up production infrastructure from scratch?
A complete production infrastructure setup — VPC, EKS/ECS cluster, RDS, Elasticache, CI/CD pipeline, observability stack and security baseline — takes 3–5 weeks. This includes Terraform code, documentation and knowledge transfer. Simple setups (single service, managed DB, ECS Fargate, GitHub Actions) can be done in 1–2 weeks. Complex multi-region, multi-cluster architectures take 6–10 weeks.
Should we use Kubernetes or ECS?
For most applications under 10 services: ECS Fargate. It is managed, simpler to operate, integrated with AWS services and removes the cluster management overhead that Kubernetes requires. For applications with many services, complex deployment requirements (canary, blue/green at scale), or teams with existing Kubernetes expertise: EKS. The Kubernetes ecosystem (Helm, ArgoCD, KEDA, etc.) is richer, but the operational overhead is real and requires dedicated attention.
What does an infrastructure audit cover?
We review: IAM policies for least-privilege violations and unused roles, security group rules for overly permissive access, publicly accessible resources (S3 buckets, RDS instances, EC2 with public IPs), secrets management (anything in plaintext that should be in Secrets Manager), monitoring coverage (what has no alerts), cost inefficiencies (over-provisioned instances, unattached volumes, data transfer patterns), and backup/DR coverage.
How do you handle zero-downtime deployments?
For stateless services: rolling updates with appropriate readiness probes, PodDisruptionBudgets in Kubernetes, or blue/green deployment via weighted target groups in ALB. For database migrations: backward-compatible migrations that can be deployed separately from application changes, using the expand-contract pattern. For stateful services: careful sequence of infrastructure changes before application changes. Zero-downtime is achievable for all service types with appropriate planning — it is not a special capability, it is a standard requirement.
What is GitOps and should we use it?
GitOps means git is the single source of truth for cluster state — desired state is declared in git, and an operator (ArgoCD or Flux) continuously reconciles actual state with desired state. Benefits: all cluster changes are reviewed via pull request, git history is the audit trail, rollback is a git revert, and drift detection is automatic. We recommend GitOps for any Kubernetes deployment with more than one engineer touching infrastructure. It is the standard for production Kubernetes.
How much should we be spending on cloud infrastructure?
Rough benchmarks: a typical SaaS application with moderate traffic should spend 10–20% of revenue on infrastructure. Startups often over-provision early (running m5.2xlarge when t3.medium would suffice) and then never right-size. Our cost optimisation engagements typically find 20–40% savings through right-sizing, Reserved Instance purchases, and architectural improvements like caching frequently-queried data and optimising data transfer patterns.
Do you support multi-cloud?
Yes, but we recommend against it unless you have a specific compliance or resilience requirement that mandates it. Active-active multi-cloud doubles infrastructure complexity and operational overhead without proportional reliability benefit for most applications. Multi-region within a single cloud (AWS ap-south-1 + us-east-1) provides most of the resilience benefit at a fraction of the complexity. We design for the resilience requirement, not for the architecture diagram.
What observability stack do you recommend?
For teams that want to move fast: Datadog. It covers metrics, logs, traces, APM and alerting in a single product with minimal setup time. Cost scales with host count — budget $15–25/host/month. For cost-sensitive or self-hosted requirements: Prometheus + Grafana (metrics), Loki (logs) and Jaeger/Tempo (traces) — the open-source stack with no per-host cost but significant setup and maintenance overhead. OpenTelemetry for instrumentation regardless of backend — it keeps your options open.