PROPELOO

BACKEND ENGINEERING / SERVER-SIDE DEVELOPMENT

Build the backend that holds under load you did not anticipate.

PROPELOO engineers production backend systems — from API design and database modelling through authentication, background jobs, caching strategy and the observability that tells you what is actually happening at scale. The backend is not the boring part. It is the part that determines whether your product works when it matters.

A backend that works for 100 users and fails for 10,000 was never production-ready — it was prototype-ready.

Most backend systems fail under load for predictable reasons: no connection pooling (database connections exhausted at 500 concurrent users), synchronous processing of operations that should be async (API response time spikes when email sending is inline), missing indexes on query patterns that only become slow at production data volume, no caching layer so the same expensive query runs thousands of times per minute. These are not exotic failure modes — they are the ones we see in every backend review of a system that was built fast and never hardened for production. PROPELOO builds backends with the production failure modes modelled from the start: connection pooling configured, async jobs extracted, query plans reviewed and caching designed before the first user arrives.

The backend engineering stack.

System Layers

  • API Layer: REST or GraphQL, request validation, authentication middleware, rate limiting, versioning
  • Business Logic Layer: Domain services, workflow orchestration, rule engines, external integrations
  • Data Layer: Database schema, query optimisation, migrations, caching, connection pooling
  • Async Layer: Background jobs, message queues, event-driven processing, scheduled tasks
  • Observability Layer: Structured logging, distributed tracing, metrics, error monitoring, alerts

Core Technical Capabilities

  • API Engineering

    REST APIs with OpenAPI spec, GraphQL with DataLoader (N+1 prevention), request validation with Zod/Joi, consistent error responses (RFC 7807) and API versioning strategy.

  • Authentication & Authorisation

    JWT with short expiry and refresh token rotation, OAuth 2.0 integration, RBAC at the service layer (not just the UI), session invalidation and audit logging for sensitive operations.

  • Database Engineering

    Schema design for access patterns, index strategy, query analysis with EXPLAIN, connection pooling via pgBouncer, read replicas for reporting queries, migrations with rollback.

  • Async Processing

    Background job queues (BullMQ, pg-boss) for email, notifications, heavy processing. Message queues (SQS, Kafka) for service communication. Scheduled tasks with distributed locking.

  • Caching Architecture

    Redis for session storage, rate limiting counters, expensive query results and computed values. Cache invalidation strategy per data type. Cache-aside vs read-through vs write-through selection.

  • Observability

    Structured JSON logging with correlation IDs, OpenTelemetry distributed tracing, Prometheus metrics endpoints, Sentry error monitoring and p95/p99 latency dashboards.

How we think about backend engineering.

The backend is a set of bets on access patterns and load characteristics. Get the bets wrong and you re-architecture under production pressure.

  • Design for access patterns, not entities

    A database schema designed around domain entities ("what data do we have?") will have different performance characteristics than one designed around access patterns ("how will we query this data?"). The indexes that matter are the ones on columns that appear in WHERE clauses of frequent queries. Design the queries first, then the schema.

    Axiom:

  • Async everything that does not need to be sync

    An API endpoint that sends an email, processes an image, calls three third-party APIs and updates five database tables while the user waits will have variable and unpredictable response times. Extract everything that does not need to happen before the response into background jobs. The endpoint returns in 50ms. The jobs process in seconds. The user experience is better.

    Axiom:

  • Caching is not optimisation — it is architecture

    A backend that makes the same expensive database query 10,000 times per minute will fail at scale regardless of hardware. Redis caching of expensive queries, computed aggregations and session state is not a performance optimisation to add later — it is an architectural decision that must be designed when the data access patterns are first understood.

    Axiom:

  • Connection pooling or connection exhaustion

    PostgreSQL has a maximum connection limit (typically 100-200). Each Node.js process without a connection pool opens a new database connection per request. At 50 concurrent requests across 10 Node.js instances, you exhaust the connection limit. pgBouncer (connection pooler) or built-in connection pool configuration is not optional for production backends.

    Axiom:

The backend architecture decisions.

  • Language and runtime?

    Impact: Node.js (TypeScript) for most API backends — ecosystem, developer availability and performance are sufficient for most workloads. Go for high-throughput microservices where latency and CPU efficiency matter. Python for backends tightly coupled to ML/data science.

    • Node.js (TypeScript) — fast to develop, large ecosystem, single-threaded (use clustering/workers for CPU)
    • Go — compiled, fast, excellent concurrency, smaller ecosystem, lower developer familiarity
    • Python (FastAPI) — excellent for ML/data, slower than Go/Node for high-throughput APIs
    • Java/Kotlin (Spring Boot) — enterprise-grade, verbose, highest hiring pool
  • REST vs GraphQL?

    Impact: REST for public-facing APIs. GraphQL when multiple clients (mobile, web, partner) need different data shapes. tRPC for TypeScript full-stack apps with shared types.

    • REST + OpenAPI — universal, excellent tooling, correct for most APIs
    • GraphQL — client-defined queries, good for complex multi-client data
    • tRPC — TypeScript end-to-end types, internal APIs only
    • gRPC — internal services, binary protocol, strong typing
  • Database primary?

    Impact: PostgreSQL for almost everything. Its JSONB support handles semi-structured data. Full-text search handles most search requirements. ACID transactions handle financial data. Only deviate with strong justification.

    • PostgreSQL — ACID, JSONB, full-text search, best general-purpose choice
    • MySQL — similar to PostgreSQL, strong if team has MySQL expertise
    • MongoDB — document model, weaker consistency, appropriate for specific use cases
    • DynamoDB — infinite scale, restrictive queries, high cost at volume
  • Background job queue?

    Impact: BullMQ for Node.js backends with Redis already in the stack. pg-boss for backends that want to avoid Redis dependency. Temporal for long-running multi-step workflows that must survive restarts.

    • BullMQ (Redis-backed) — feature-rich, retries, priorities, delayed jobs, Node.js native
    • pg-boss (PostgreSQL-backed) — leverages existing PostgreSQL, no Redis required
    • Temporal — durable workflows, long-running processes, higher setup complexity
    • AWS SQS — managed, unlimited scale, no persistence concern
  • Authentication approach?

    Impact: Managed auth (Auth0, Clerk) for speed to market. JWT with refresh rotation for teams that want control. Session cookies for internal tools where stateless auth overhead is not worth it.

    • JWT access + refresh tokens — stateless, scalable, token rotation required
    • Session cookies — stateful, requires session store (Redis), simpler revocation
    • Auth0/Clerk (managed auth) — fastest implementation, ongoing cost
    • Custom OAuth 2.0 — maximum control, significant implementation risk
  • Caching strategy?

    Impact: Redis cache-aside for session storage, rate limiting and expensive query results. CDN caching for public API responses with predictable staleness tolerance.

    • No cache — simplest, fails at scale
    • Redis cache-aside — application reads cache first, populates on miss
    • Redis read-through — transparent caching layer
    • CDN caching — for public, cacheable API responses

What PROPELOO builds.

  • REST API Backend

    Node.js/Go REST API with OpenAPI spec, JWT auth, PostgreSQL, Redis caching, background jobs and production monitoring.

  • GraphQL API

    Apollo Server or Pothos GraphQL with DataLoader, subscriptions, persisted queries and complexity limiting.

  • Event-driven Backend

    Kafka or SQS-based event-driven architecture — producers, consumers, dead letter queues, retry logic and event schema registry.

  • High-throughput API

    Go backend engineered for 50K+ requests/second — connection pooling, read replicas, Redis caching and horizontal scaling.

  • Backend Audit & Optimisation

    Review of existing backend for N+1 queries, missing indexes, connection pool exhaustion, sync operations that should be async and missing observability.

  • API Platform with Developer Portal

    Public API backend with OAuth 2.0, API key management, rate limiting, OpenAPI documentation, SDK generation and developer sandbox.

The backend stack.

  • Runtime

    Stack: Node.js (TypeScript), Go, Python (FastAPI), Bun

  • API

    Stack: Fastify, Express, Apollo GraphQL, tRPC, Hono

  • Database

    Stack: PostgreSQL, Redis, Prisma / Drizzle ORM, pgBouncer, pg-boss

  • Queue & Events

    Stack: BullMQ, Temporal, Kafka, AWS SQS, RabbitMQ

  • Auth

    Stack: JWT (jose), Auth0, Clerk, Passport.js, Custom OAuth 2.0

  • Observability

    Stack: OpenTelemetry, Sentry, Datadog, Prometheus, Pino (logging)

Backend security prevents the most common production breaches.

  • SQL Injection Prevention

    Parameterised queries — never string concatenation for SQL. ORM query builders (Prisma, Drizzle, Knex) use parameterised queries by default. Raw SQL must be reviewed carefully.

  • Input Validation

    Validate all request inputs at the API boundary using Zod or Joi — string length, format, type, enum values. Reject unknown fields. Never trust client-provided data.

  • Rate Limiting

    Per-API-key or per-IP rate limiting using Redis counters. Express-rate-limit, Fastify rate-limit or custom Redis middleware. Prevents brute force and DDoS without infrastructure changes.

  • Secrets Management

    Environment variables from AWS Secrets Manager or HashiCorp Vault at runtime. No credentials in source code. Pre-commit hooks scan for accidental secrets.

  • CORS Configuration

    Explicit allow-list for CORS origins — not wildcard (*) for credential-bearing requests. Preflight caching headers for performance. Separate CORS config for public and authenticated endpoints.

  • Audit Logging

    Structured log entries for all authentication events (login, logout, token refresh, failed attempts), sensitive data access and admin operations. Logs to immutable store for security review.

From design to production backend.

  1. 01. API Design

    OpenAPI spec written before implementation. Resource modelling, auth design, error taxonomy.

  2. 02. Foundation

    Project scaffold, database schema, migrations, auth middleware, CI pipeline.

  3. 03. Core Endpoints

    Business logic endpoints with validation, error handling and unit tests.

  4. 04. Async Layer

    Background job queue, scheduled tasks, email/notification workers.

  5. 05. Caching & Performance

    Redis integration, query optimisation with EXPLAIN, connection pool tuning.

  6. 06. Observability

    Structured logging, Sentry errors, OpenTelemetry tracing, latency dashboards.

  7. 07. Load Testing & Launch

    k6 load test to expected peak, fix bottlenecks, production deployment.

Frequently Asked Questions

Node.js vs Go for backend?

Node.js for most product backends — large ecosystem (npm), TypeScript support, fast development, sufficient performance for most web APIs (handles 10K req/s on a single core with async I/O). Go for high-throughput services where latency and CPU efficiency matter — Go handles 50K-100K req/s on comparable hardware, with lower memory usage and better concurrency primitives. Choose Go when performance requirements are demonstrated, not anticipated.

How do we prevent N+1 database queries?

N+1 queries occur when fetching a list (1 query) then fetching related data per item (N queries). Solutions: DataLoader for batching in GraphQL, eager loading with JOIN or ORM includes for REST, query analysis with pg_stat_statements in PostgreSQL. Every GraphQL resolver should use DataLoader. Every ORM query returning a list should specify which relations to include.

When do we need a message queue?

Message queues are justified when: processing must continue if the consumer is temporarily unavailable, processing is too slow for synchronous response (image processing, PDF generation), multiple services need to react to the same event, or traffic spikes must be absorbed without losing work. Common mistake: using message queues prematurely when a simple background job queue (BullMQ) would suffice.

How do we handle database migrations safely?

Expand-contract pattern for zero-downtime migrations: expand (add new column/table without removing old), migrate application to use new structure, contract (remove old column/table after all instances use new structure). Never drop columns in the same migration as adding them — old application instances still reference them. Automated migration in CI (run on deploy), rollback scripts prepared and tested.