PROPELOO

AI API INTEGRATION / LLM INTEGRATION

Integrate AI APIs that actually work reliably in production.

PROPELOO engineers AI API integrations — from OpenAI, Anthropic and Google to open-source model APIs — with streaming, fallback routing, cost controls, rate limit handling, structured output validation and the monitoring that tells you when your AI integration is degrading. Integrating an AI API is one hour. Building a production AI API integration is weeks.

An AI API that works in development will fail in production in ways that are specific to AI.

AI API production failures are different from standard API failures. They are not just downtime and rate limits — they are hallucinations that reach users, latency spikes when the provider's servers are overloaded, silent quality degradations when the provider updates the model, and cost spikes when a single request triggers an unexpectedly long generation. PROPELOO builds AI API integrations with streaming for UX, fallback routing across providers for reliability, structured output validation for correctness, cost controls for budget, and monitoring for quality — because AI APIs require their own category of production engineering.

Production AI API integration stack.

System Layers

  • API Layer: Provider SDKs, authentication, model selection, parameter configuration
  • Streaming Layer: Server-sent events, token streaming, partial result handling
  • Reliability Layer: Retry logic, fallback providers, timeout handling, circuit breakers
  • Validation Layer: Structured output parsing, Pydantic validation, output type enforcement
  • Observability Layer: Cost tracking, latency monitoring, quality metrics, provider status

Core Technical Capabilities

  • Multi-provider Integration

    OpenAI, Anthropic, Google Gemini, Cohere and Groq integrated behind a single interface — route to the best model per task, fall back when a provider is degraded, A/B test between providers.

  • Streaming Implementation

    Server-sent events (SSE) for streaming responses to web/mobile clients. Partial result accumulation, streaming to database, client reconnection handling and streaming through API gateways.

  • Structured Output

    OpenAI function calling / JSON mode, Anthropic tool use, Pydantic validation schemas with retry on validation failure. Transform unstructured LLM output into typed, validated application data.

  • Cost Control

    Per-request token counting, per-user and per-day spend limits, cost attribution by feature/user/team, spend alerts and automatic model downgrade when budget thresholds are approached.

  • Fallback & Reliability

    Automatic fallback: if OpenAI is degraded, route to Anthropic. Circuit breaker per provider. Retry with exponential backoff for transient errors. SLA monitoring per provider with alerting.

  • AI Proxy Layer

    Centralised LiteLLM proxy or custom gateway — API key management, unified interface across providers, request logging, rate limiting per downstream client, cost allocation.

How we think about AI API integration.

AI API integration is not different from any other external API integration — except for the hallucination risk, the streaming requirement and the 10-100x cost variance between models.

  • Provider abstraction saves future pain

    Applications that call the OpenAI SDK directly throughout the codebase are locked to OpenAI. A thin abstraction layer (LiteLLM or custom) that routes to any provider means: switch providers in one place, A/B test providers without code changes, fall back automatically on degradation. Build the abstraction from day one.

    Axiom:

  • Streaming is a UX requirement, not a performance optimisation

    An LLM that generates 500 tokens takes 5-10 seconds. Displayed all at once after generation: perceived as slow. Streamed token-by-token: perceived as fast. Streaming is a 1-2 day implementation decision with significant UX impact.

    Axiom:

  • Validate structure, monitor quality

    Structured output validation (Pydantic schemas + retry on failure) catches format errors. LLM-as-judge or RAGAS metrics catch quality degradation when a provider silently updates a model. Both are required in production.

    Axiom:

  • Cost spikes are infrastructure incidents

    A single prompt that generates 50,000 tokens costs $0.25 at GPT-4o pricing. At 1,000 requests/day: $250/day from one prompt pattern. Per-request token limits, prompt length validation and cost monitoring with alerts are required production safeguards.

    Axiom:

AI API integration decisions.

  • Direct SDK vs LiteLLM proxy?

    Impact: LiteLLM for any application that might ever switch providers or needs multi-provider fallback. Direct SDK only for simple single-provider integrations.

    • Direct OpenAI/Anthropic SDK — simplest, provider-locked
    • LiteLLM — unified interface, 100+ providers, self-hosted or cloud
    • Custom proxy — full control, significant maintenance
    • Portkey/Helicone — managed proxy, observability features
  • Streaming approach?

    Impact: SSE from backend API route to client for all streaming LLM responses. Never expose provider API keys to client-side code.

    • SSE from backend to client — standard, works everywhere
    • Direct client-to-provider streaming — simpler, exposes API key
    • Websocket streaming — bidirectional, overkill for most cases
    • Batch (no streaming) — simplest, poor UX for long generations
  • Structured output method?

    Impact: Instructor (Python) for provider-agnostic structured output with Pydantic validation and automatic retry on validation failure. Most reliable approach across providers.

    • OpenAI function calling / JSON mode — reliable, GPT-specific
    • Instructor library — provider-agnostic, Pydantic integration
    • Prompt engineering only — unreliable for complex schemas
    • Regex/parsing post-processing — fragile
  • Cost control strategy?

    Impact: Layered: per-request token limits (prevent runaway single requests) + per-user daily limits (prevent abuse) + spend alerts (operational visibility).

    • No controls — discover costs after the fact
    • Per-request token limits — cap individual request cost
    • Per-user daily limits — cap per-user spend
    • Budget alerts + model downgrade — alert then switch to cheaper model
  • Provider fallback?

    Impact: Automatic fallback to a different provider (OpenAI → Anthropic) provides real resilience. Same-provider multi-region helps with rate limits but not provider-wide outages.

    • No fallback — single provider, single point of failure
    • Manual fallback (change config) — requires human intervention
    • Automatic fallback — route to backup on error
    • Multi-region same provider — redundancy within provider
  • Caching strategy?

    Impact: Exact match cache for deterministic prompts (classification, extraction with identical inputs). Semantic cache for FAQ-style applications. No cache for creative generation.

    • No cache — every request hits LLM API
    • Exact match cache — same prompt = cached response
    • Semantic cache (GPTCache) — similar prompts share cache
    • Deterministic output cache — temperature=0, high hit rate

What PROPELOO integrates.

  • LLM-powered Product Feature

    Integrate OpenAI/Anthropic into an existing product — streaming API route, cost controls, structured output and quality monitoring.

  • Multi-provider AI Platform

    LiteLLM proxy with provider fallback, cost attribution by team/feature, usage dashboards and automatic model routing by task.

  • Document Processing Pipeline

    Batch AI processing pipeline — structured extraction from documents, validation, retry on failure, cost-efficient model selection per task complexity.

  • AI Feature Backend

    FastAPI backend for AI features — streaming endpoints, authentication, rate limiting, cost tracking and provider abstraction.

  • Legacy App + AI

    Add AI capabilities to existing application — API integration that does not require rewriting the existing stack, streaming frontend components.

  • AI API Cost Audit

    Audit existing AI API usage — identify cost inefficiencies, over-sized models for tasks, missing caching and optimise for 50-70% cost reduction.

The AI API integration stack.

  • Provider SDKs

    Stack: openai (Python/Node), anthropic (Python/Node), google-generativeai, cohere, groq

  • Abstraction

    Stack: LiteLLM, Portkey, Helicone, Custom proxy (FastAPI)

  • Structured Output

    Stack: Instructor (Python), Pydantic v2, OpenAI JSON mode, Zod (TypeScript)

  • Streaming

    Stack: Server-sent events (SSE), FastAPI StreamingResponse, Vercel AI SDK (Next.js)

  • Monitoring

    Stack: LangSmith, Helicone, Datadog, Custom token counter

  • Backend

    Stack: FastAPI, Node.js (Fastify), Redis (cache), PostgreSQL (logs)

AI API security in production.

  • API key management

    Provider API keys server-side only. Never in frontend bundles or environment variables committed to git. AWS Secrets Manager or equivalent. Per-environment separate keys.

  • Spend controls

    OpenAI and Anthropic both offer hard spending limits via the provider dashboard. Set these as a backup to application-level controls. A leaked API key with no spend limit can cost thousands in hours.

  • Rate limiting per user

    Per-user request rate limiting prevents abuse. Without rate limiting, a single determined user can exhaust your AI API budget. Redis-based rate limiting per user/API key.

  • Output sanitisation

    AI-generated content rendered in HTML must be sanitised — LLMs can generate XSS vectors in certain prompts. Treat all LLM output as untrusted user input before rendering.

  • PII in prompts

    User data sent to third-party LLM APIs may violate GDPR/HIPAA. PII detection and redaction before API calls, data processing agreements with providers, self-hosted models for regulated data.

  • Prompt injection

    User input included in prompts can override system instructions. Input sanitisation, separate user and system context, and LLM-based injection detection for high-risk applications.

From API key to production integration.

  1. 01. Integration Architecture

    Provider selection, abstraction approach, streaming design, cost model.

  2. 02. API Layer

    Provider SDK integration, abstraction layer, authentication, error handling.

  3. 03. Streaming

    SSE endpoint, client streaming component, reconnection handling.

  4. 04. Structured Output

    Schema definition, Instructor/function calling, validation and retry.

  5. 05. Cost Controls

    Token counting, per-user limits, spend alerts, model routing.

  6. 06. Fallback & Reliability

    Multi-provider routing, circuit breakers, retry logic.

  7. 07. Monitoring

    Cost dashboard, latency tracking, quality monitoring, provider status alerts.

Frequently Asked Questions

What is LiteLLM?

LiteLLM is an open-source proxy that provides a unified OpenAI-compatible API interface to 100+ LLM providers (OpenAI, Anthropic, Google, Cohere, Bedrock, etc.). Your application calls the LiteLLM interface; LiteLLM routes to the configured provider. Benefits: switch providers without code changes, A/B test between models, automatic fallback on provider errors, unified cost tracking.

How do we handle AI API outages?

Automatic fallback: configure primary and secondary providers. If OpenAI returns 5xx errors, route to Anthropic. LiteLLM handles this natively. Circuit breaker: after N consecutive failures from a provider, mark it as degraded for T minutes and route all traffic to the fallback. Monitor provider status pages (OpenAI, Anthropic both have status pages) and alert on degradation.

How do we control AI API costs?

Four layers: (1) Per-request max tokens (cap individual response length). (2) Per-user daily token budget with Redis counter. (3) Use smaller/cheaper models for simple tasks (GPT-4o-mini for classification, GPT-4o for complex reasoning). (4) Cache responses for identical or semantically similar prompts. Together these can reduce costs by 50-80% vs unoptimised usage.