PROPELOO

MICROSERVICES / DISTRIBUTED SYSTEMS

Build microservices when the problem actually requires them.

PROPELOO engineers microservices architectures — from service decomposition and API contract design through event-driven communication, service mesh, distributed tracing and the operational infrastructure that makes distributed systems maintainable. Microservices solve specific problems. They also introduce specific problems. We start with whether microservices are the right answer before designing the architecture.

Most systems that are built as microservices should have started as a monolith.

Microservices solve three specific problems: independent scaling of different components, independent deployment by different teams, and technology heterogeneity where different services genuinely need different stacks. If your system does not have these requirements, microservices introduce network latency, distributed transaction complexity, operational overhead and debugging difficulty without proportional benefit. A well-structured monolith ships faster, is easier to test, cheaper to run and simpler to debug than a microservices architecture at the same feature scope. PROPELOO helps you make the correct architectural decision — and when microservices are the right answer, builds them correctly.

The full microservices engineering stack.

Microservices architecture requires engineering decisions across six domains simultaneously.

System Layers

  • Service Design Layer: Domain-driven decomposition, service boundaries, API contracts, data ownership
  • Communication Layer: Synchronous (REST/gRPC), asynchronous (Kafka/SQS), event schemas, contract testing
  • Data Layer: Database per service, eventual consistency, saga pattern, event sourcing
  • Infrastructure Layer: Container orchestration, service discovery, load balancing, health checks
  • Observability Layer: Distributed tracing, centralised logging, metrics, alerting, correlation IDs

Core Technical Capabilities

  • Service Decomposition

    Domain-driven design for service boundary identification — bounded contexts, aggregate roots, domain events. Avoiding the most common mistake: decomposing by technical layer (database service, API service) rather than business capability.

  • API Design & Contracts

    OpenAPI for REST, Protobuf for gRPC, AsyncAPI for event schemas. Consumer-driven contract testing with Pact. API gateway for external-facing services. Backend-for-frontend (BFF) pattern for client-specific API aggregation.

  • Event-driven Architecture

    Kafka or AWS SQS/SNS for async service communication. Event sourcing for audit trail and replay capability. Saga pattern for distributed transactions. Outbox pattern for reliable event publishing without two-phase commit.

  • Data Management

    Database-per-service pattern, polyglot persistence (right database per service), CQRS for read/write separation, eventual consistency strategies and data consistency patterns across service boundaries.

  • Service Mesh

    Istio or Linkerd for service-to-service mTLS, traffic management, circuit breaking, retry logic and observability at the infrastructure layer — without code changes in individual services.

  • Distributed Tracing

    OpenTelemetry instrumentation across all services, Jaeger or Zipkin for trace visualisation, correlation ID propagation and service dependency mapping for performance analysis and incident debugging.

How we think about microservices.

Microservices are an organisational solution as much as a technical one. Conway's Law: systems reflect the communication structure of the organisation that builds them. If you do not have the team structure for microservices, you will not get the benefits.

  • Monolith first, extract on pain

    Start with a well-structured modular monolith. Extract services when you have a concrete, demonstrated need: a specific component that needs to scale independently, a team that needs independent deployment velocity, or a service that genuinely benefits from a different technology stack. Extracting services based on architectural ideals rather than demonstrated need is premature optimisation at the system level.

    Axiom:

  • Service boundaries must align with data ownership

    A service that queries another service's database is not a microservice — it is a distributed monolith. Each service must own its data exclusively. If two services need the same data, either they belong in the same service or one needs to maintain its own copy via events. Wrong service boundaries are the most common cause of distributed systems complexity that exceeds the complexity of the monolith it replaced.

    Axiom:

  • Distributed transactions are not free

    A transaction that spans two microservices cannot use database ACID guarantees. It must use a saga (sequence of local transactions with compensating transactions on failure) or accept eventual consistency. Both require significantly more engineering than a single database transaction. Design your service boundaries to minimise cross-service transactions — the fewer you need, the simpler the system.

    Axiom:

  • Observability is non-negotiable

    Debugging a monolith: add a print statement, reproduce the issue. Debugging a microservices system without distributed tracing: follow a request across 6 services, each with its own log format and timezone, trying to reconstruct a timeline. OpenTelemetry distributed tracing is not optional infrastructure for a production microservices system. It must be in place before going to production.

    Axiom:

The distributed systems decisions that matter.

Each choice propagates across every service in the system.

  • Microservices vs modular monolith?

    Impact: Default to modular monolith. Extract services when you can name the specific team or scaling requirement that justifies the operational overhead. Most successful startups ran as monoliths through their high-growth phase.

    • Microservices — independent deployment, independent scaling, operational complexity
    • Modular monolith — clean module boundaries, single deployment, simpler operations
    • Mini-services (2-5 services) — pragmatic middle ground for most teams
    • Serverless functions — extreme decomposition, event-driven, cold start issues
  • Synchronous vs async communication?

    Impact: Hybrid is correct: synchronous REST/gRPC for queries (read operations that need immediate response), async events for commands (state changes that can be processed eventually). Event-driven everywhere is powerful but makes request tracing significantly harder.

    • REST everywhere — simple, universal, tight coupling
    • gRPC for internal, REST for external — typed, fast internal; flexible external
    • Event-driven everywhere — loose coupling, eventual consistency, debugging complexity
    • Hybrid — sync for queries, async for commands
  • Service discovery?

    Impact: Kubernetes DNS for K8s deployments (zero additional infrastructure). Consul for hybrid or non-K8s deployments. Never hardcode service URLs in config — service instances change too frequently.

    • Kubernetes DNS (kube-dns) — built-in, sufficient for K8s deployments
    • Consul — language-agnostic, health checking, works outside K8s
    • AWS Cloud Map — AWS-native, integrates with ECS/EKS
    • Hardcoded URLs with env vars — simple, does not scale to many services
  • Distributed transaction handling?

    Impact: Design service boundaries to avoid distributed transactions first. When unavoidable: orchestration saga with a dedicated saga orchestrator service. Temporal.io provides excellent saga orchestration tooling.

    • Choreography saga — events drive compensation, no central coordinator, harder to track
    • Orchestration saga — central coordinator, easier to understand, single point of failure
    • Two-phase commit — strong consistency, performance impact, not practical across services
    • Avoid distributed transactions — design service boundaries to eliminate the need
  • API gateway?

    Impact: Kong or AWS API Gateway for most production systems. Traefik as K8s ingress for services that do not need complex gateway features (auth, rate limiting, transformation). API gateway handles cross-cutting concerns that should not be in individual services.

    • Kong — open source, plugin ecosystem, self-hosted
    • AWS API Gateway — managed, AWS-native, pricing scales with requests
    • Nginx/Traefik — lightweight, K8s-native ingress, limited features
    • Custom gateway — maximum control, maintenance burden
  • Service mesh (yes/no)?

    Impact: Service mesh is justified for systems with >10 services where mTLS between services, traffic management and infrastructure-level observability are required. For smaller systems, the Istio complexity is not worth it — handle resilience (retries, circuit breaking) in the API client layer.

    • No service mesh — simpler, each service handles its own resilience
    • Istio — feature-rich, significant complexity, resource overhead
    • Linkerd — lighter than Istio, simpler, less features
    • AWS App Mesh — managed, AWS-native, less flexible

What PROPELOO builds.

  • Microservices Migration

    Strangler fig migration from monolith to microservices — extract high-value services first, maintain monolith during transition, validate extracted services before decommissioning monolith code.

  • Event-driven Platform

    Kafka-based event-driven architecture for high-throughput data processing — event schema registry, consumer groups, dead letter queues, replay capability and monitoring.

  • API Platform

    Internal API platform for multiple frontend clients — API gateway, BFF pattern per client, service registry, rate limiting, auth middleware and developer documentation.

  • CQRS + Event Sourcing

    Write model (commands) and read model (queries) separated with event store — complete audit trail, temporal queries and read model rebuilding from event history.

  • Multi-region Service Architecture

    Globally distributed services with data sovereignty, region-specific deployments, cross-region event replication and region failover without data loss.

  • Microservices Architecture Review

    Assessment of existing microservices architecture — service boundary analysis, data ownership violations, distributed transaction identification, observability gaps and remediation roadmap.

The microservices stack.

Communication, data, orchestration and observability each require specific tooling.

  • Services

    Stack: Node.js (TypeScript), Go, Java Spring Boot, Python (FastAPI), Rust (Axum)

  • Communication

    Stack: REST + OpenAPI, gRPC + Protobuf, Kafka, AWS SQS/SNS, NATS

  • Data

    Stack: PostgreSQL, MongoDB, Redis, Elasticsearch, DynamoDB

  • Infrastructure

    Stack: Kubernetes (EKS), Istio / Linkerd, Kong API Gateway, Terraform, ArgoCD

  • Observability

    Stack: OpenTelemetry, Jaeger / Tempo, Prometheus + Grafana, Loki, Datadog

  • Workflow

    Stack: Temporal.io, AWS Step Functions, Apache Airflow, Kafka Streams

Distributed systems have distributed attack surfaces.

Each service boundary is a potential security boundary that must be enforced.

  • Service-to-service Authentication

    Services must authenticate each other — not just authenticate end users. mTLS via service mesh (Istio) or JWT service tokens with short expiry. A compromised internal service should not have unauthenticated access to all other services.

  • API Gateway Security

    Authentication, rate limiting, request validation and WAF at the API gateway layer — before requests reach individual services. Services behind the gateway should only be reachable from within the cluster.

  • Secrets Management

    Each service has its own database credentials, API keys and certificates. HashiCorp Vault or AWS Secrets Manager with automatic rotation. Secrets injected at runtime, not baked into container images.

  • Network Segmentation

    Services should only have network access to services they need to communicate with — not all services in the cluster. Kubernetes NetworkPolicies or Istio authorization policies enforce service-to-service communication rules.

  • Distributed Tracing & Audit

    Correlation IDs propagated across all service calls enable tracing of a user request across the entire system. Security events must include the full request chain for forensic analysis.

  • Container Security

    Each service container should run as non-root with minimal base image. Trivy scanning in CI, OPA policies enforced at Kubernetes admission, no privileged containers and read-only root filesystem where possible.

From monolith to production microservices.

  1. 01. Architecture Assessment

    Evaluate whether microservices are justified. Define service boundaries via domain-driven design. Identify data ownership and communication patterns.

  2. 02. Foundation Infrastructure

    Kubernetes cluster, service mesh (if justified), API gateway, distributed tracing and CI/CD pipeline for multi-service deployments.

  3. 03. Service Scaffolding

    Service template with health checks, tracing instrumentation, structured logging, metrics and deployment manifests. Every service starts from this template.

  4. 04. Service Development

    Service-by-service development with contract testing, integration testing and independent deployment validation.

  5. 05. Event Infrastructure

    Kafka/SQS setup, event schema registry, consumer groups, dead letter queues and replay capability.

  6. 06. Observability

    Distributed tracing across all services, service dependency map, latency dashboards and cross-service alerting.

  7. 07. Runbooks & Handoff

    Per-service runbooks, on-call training, incident response procedures for distributed system failures and knowledge transfer.

Frequently Asked Questions

Should we use microservices?

Ask these questions first: Do different components need to scale independently and is that requirement demonstrated? Do different teams need to deploy independently and do those teams exist? Do different services genuinely need different technology stacks? If the answer to all three is no, a modular monolith is the correct architecture. Microservices are justified when the operational overhead is worth the independence they provide — and that is rarely true at early stage.

What is domain-driven design and why does it matter for microservices?

DDD provides a framework for identifying service boundaries — specifically the "bounded context" concept, which defines the scope within which a particular domain model is consistent. A bounded context maps naturally to a microservice boundary. Getting the boundaries wrong is the most common and most expensive microservices mistake: too fine-grained (nano-services) creates excessive inter-service communication; too coarse-grained (distributed monolith) creates tight coupling with distributed system complexity.

What is the outbox pattern?

The outbox pattern solves the dual-write problem: when a service needs to update its database AND publish an event, these two operations cannot be in the same ACID transaction across different systems. The outbox pattern writes both the state change and the event to the same database in a single transaction (the event goes to an "outbox" table). A separate process reads the outbox and publishes events to Kafka/SQS. This guarantees that events are published exactly when the state change is committed.

How do we debug issues across multiple services?

Distributed tracing with correlation IDs is the only practical answer. Every request gets a unique trace ID that is propagated across all service calls via HTTP headers or message headers. OpenTelemetry instruments your services automatically. Jaeger or Grafana Tempo visualises the complete request trace — showing which services were called, in what order, with what latency at each hop. Without distributed tracing, debugging cross-service issues relies on log correlation by timestamp, which is tedious and error-prone.