PROPELOO

AI AGENTS / AUTONOMOUS SYSTEMS

AI Agent Development — autonomous systems that act, not just answer.

PROPELOO engineers production AI agents — from task decomposition architecture and tool integration to multi-agent orchestration, safety guardrails and the evaluation infrastructure that makes agents measurable. An agent demo takes a day to build. A production agent takes a different kind of engineering.

An agent that works 90% of the time is not a product. It is a liability.

Agent reliability is architectural. Task decomposition strategy, tool call reliability, context window management, error handling and recovery, human-in-the-loop design, evaluation at scale — each a separate engineering problem. Most agent demos work in the happy path. Production agents need edge case handling, graceful retry logic, deterministic escalation to humans, and evaluation frameworks that catch regressions before users do. PROPELOO builds agents designed for the 10% of cases that break everything.

What sits underneath a production AI agent.

An AI agent is not an LLM with tool calling enabled. It is a distributed system with task planning, tool dependencies, state management and failure modes at every step.

System Layers

  • Task Planning Layer: Goal decomposition, subtask generation, dependency ordering, plan validation, replanning on failure
  • Tool Integration Layer: Tool definitions, API integrations, input validation, output parsing, error handling per tool
  • Memory & Context Layer: Short-term context management, long-term memory store, semantic retrieval, context window optimisation
  • Execution & Safety Layer: Action execution, sandboxing, rate limiting, human approval gates, rollback capabilities, audit logging
  • Evaluation & Monitoring: Task completion rate, tool call accuracy, latency tracking, cost monitoring, regression test suites

Core Technical Capabilities

  • Agent Architecture Design

    Full agent system design — reasoning strategy selection (ReAct, Plan-Execute, custom), state machine specification, tool dependency mapping and failure mode analysis before any code is written.

  • Multi-agent Orchestration

    Supervisor-worker architectures, specialised agent networks, inter-agent communication protocols, shared memory systems and coordination patterns for complex multi-step workflows.

  • Tool & API Integration Layer

    Robust tool definitions with strict input/output schemas, error handling for all failure modes, retry logic, timeout management and safe execution boundaries for every external API call.

  • Memory & Context Management

    Short-term context window management with intelligent pruning, long-term memory using vector stores, episodic memory for task history and working memory for multi-session continuity.

  • Safety Guardrails System

    Input and output filtering, action scope restrictions, human-in-the-loop approval gates for irreversible actions, anomaly detection for unusual agent behaviour and automatic escalation paths.

  • Agent Evaluation & Testing

    End-to-end task evaluation frameworks measuring completion rate, tool call accuracy, path efficiency, edge case handling and regression performance across model updates and prompt changes.

How we think about AI agent engineering.

An AI agent is only as reliable as the system you build around it — the task planner, tool definitions, error handling and human escalation path.

  • Agents are systems, not features

    An AI agent is a distributed system with tool dependencies, state management, retry logic and failure modes. It requires the same reliability engineering as any microservice — circuit breakers, timeouts, dead letter queues and runbooks.

    Axiom:

  • Happy path demos are not products

    Most agent demos run the optimal path with clean inputs and available tools. Production agents face ambiguous inputs, flaky APIs, context window limits and tasks the original design did not anticipate. Every failure mode must have a defined behaviour.

    Axiom:

  • Scope determines safety

    Every action an agent can take autonomously is a surface area for error. The boundary between what the agent does autonomously and what requires human approval must be explicit, documented and implemented in code — not in a prompt.

    Axiom:

  • Evaluation is the safety net

    Without continuous evaluation, agent quality degrades invisibly after model updates or tool changes. Task completion rate, path efficiency and edge case handling must be measured continuously — not checked manually before releases.

    Axiom:

The decisions that define an AI agent system.

These architectural choices determine reliability, cost and the surface area for catastrophic failures.

  • Single agent vs multi-agent architecture

    Impact: Multi-agent architectures multiply capability but also multiply failure modes. Each agent boundary is a potential communication failure, state inconsistency and debugging nightmare.

    • Single agent — simpler state management, lower coordination overhead, limited parallel task execution
    • Supervisor + worker agents — parallel task execution, role specialisation, more complex debugging
    • Peer-to-peer multi-agent — maximum flexibility, highest coordination complexity and failure surface
  • ReAct vs Plan-Execute vs custom reasoning

    Impact: Reasoning strategy determines how the agent handles unexpected states. ReAct is natural but prone to getting stuck in loops. Plan-Execute is predictable but cannot adapt to mid-task discoveries.

    • ReAct (Reason + Act) — interleaved reasoning and action, natural for exploratory tasks, can spiral on complex problems
    • Plan-Execute — generate full plan then execute, more predictable, struggles with plan-invalidating discoveries
    • Custom orchestration — fully controlled execution flow, highest reliability, highest engineering cost
  • Stateful vs stateless agent sessions

    Impact: Stateful agents can resume interrupted tasks and learn from past runs. They also require careful state invalidation logic and create dependency on the state store availability.

    • Stateless — each run independent, simple infrastructure, no cross-session memory
    • Stateful with DB persistence — cross-session memory, task resumability, state management complexity
    • Hybrid — stateless compute with persistent memory store
  • Synchronous vs asynchronous agent execution

    Impact: Tasks taking more than 30 seconds should run asynchronously. Synchronous long-running agents create HTTP timeout failures and poor user experience.

    • Synchronous — blocking execution, simple client model, poor UX for long tasks
    • Async with job queue — non-blocking, better for long-running tasks, requires job management infrastructure
    • Streaming with real-time updates — best UX for interactive agents, complex infrastructure
  • Human-in-the-loop vs fully autonomous

    Impact: Every irreversible action an agent takes without human approval is a liability. The tier of autonomy must match the reversibility and cost of each action type explicitly.

    • Fully autonomous — maximum efficiency, highest risk for irreversible actions
    • Approval gates for high-stakes actions — good balance, requires defining "high-stakes" explicitly
    • Human review for all output — maximum safety, defeats the purpose of automation
    • Tiered autonomy — autonomous for read, approval for write, human review for delete/send
  • OSS framework vs custom orchestration

    Impact: Framework abstractions reduce initial build time but create debugging complexity and version dependency risk. Complex production agents often require migrating away from frameworks as requirements evolve.

    • LangChain — extensive integrations, high abstraction, version instability, debugging complexity
    • LangGraph — stateful graph-based workflows, better for complex agent topologies
    • AutoGen / CrewAI — multi-agent specialisation, opinionated architecture
    • Custom orchestration — full control, no abstraction overhead, highest engineering investment

What AI agent engineering becomes.

  • Autonomous Research Agent

    Web research, document synthesis, competitive analysis and report generation — multi-step agent with search tools, document readers, structured output and human review gates.

  • Customer Support Agent

    Tier-1 support automation with CRM integration, knowledge base retrieval, ticket creation, escalation logic and satisfaction measurement — designed to handle the cases it can and know when it cannot.

  • Code Review Agent

    Automated pull request review — security vulnerability detection, style compliance, test coverage analysis and documentation quality assessment with inline comment generation.

  • Sales Outreach Automation

    Lead research, personalised outreach generation, CRM update automation and follow-up scheduling — with human approval gates for outbound sends and quality scoring.

  • Data Analysis Agent

    Autonomous data analyst that executes SQL queries, generates visualisations, interprets results and produces structured reports from natural language task descriptions.

  • Document Processing Agent

    Intake, classification, extraction and routing of unstructured documents — contracts, invoices, applications — with structured output, confidence scoring and human review queues.

The AI agent engineering stack.

Framework selection follows task complexity, reliability requirements and team size — not recency.

  • Orchestration Frameworks

    Stack: LangGraph, LangChain, AutoGen, CrewAI, Custom Orchestration, Temporal (workflows)

  • Foundation Models

    Stack: OpenAI GPT-4o, Anthropic Claude 3.5, Google Gemini, Mistral, Llama 3.1 (self-hosted), AWS Bedrock

  • Tool & API Integration

    Stack: OpenAPI / Swagger auto-gen, Zapier / Make connectors, Browser automation (Playwright), Code execution (E2B), Search APIs (Tavily, Serper), Custom REST tools

  • Memory & Storage

    Stack: Pinecone (vector memory), pgvector, Redis (working memory), PostgreSQL (task history), Zep (conversation memory)

  • Evaluation & Monitoring

    Stack: LangSmith, Arize Phoenix, Custom eval pipelines, Datadog, Sentry (error tracking), Weights & Biases

  • Infrastructure

    Stack: Python / FastAPI, AWS Lambda / ECS, Celery / Redis Queue, Docker, Kubernetes, Temporal (durable workflows)

AI agents introduce unique security risks beyond standard application security.

An agent with tool access and partial autonomy is a new category of attack surface. Standard security reviews do not cover it.

  • Prompt Injection & Jailbreak

    Malicious content in tool outputs or retrieved documents can hijack agent instructions mid-task — causing the agent to perform attacker-controlled actions rather than intended ones. This is indirect prompt injection. Every tool output must be treated as untrusted input, not trusted system context.

  • Tool Misuse & Unintended API Calls

    An agent with access to write APIs (email, CRM, database) can be manipulated into performing unintended actions at scale — bulk deletion, mass message sends, data exfiltration. Strict tool scope definitions, action sandboxing and rate limits per tool are required.

  • Data Leakage Through Tool Outputs

    Agents retrieving from RAG systems or databases can surface data the user should not access if access control is not enforced at the retrieval layer. Tool permission models must respect the same authorisation boundaries as direct user API access.

  • Autonomous Deletion & Spending Risk

    Agents with delete or spend permissions can cause irreversible harm at machine speed. Human-in-the-loop approval gates are required for any action that deletes data, sends external communications or incurs financial cost above a defined threshold.

  • PII in Agent Context

    Sensitive personal data surfaced through tool calls enters the agent's context window and may be sent to the LLM provider. PII must be identified and redacted before entering the context, and tool definitions must explicitly scope what data can be retrieved.

  • Agent Chain Attacks

    In multi-agent systems, a compromised agent can poison the inputs of downstream agents. Each agent must validate its inputs independently — never trust that upstream agents produced safe outputs. Agent boundary validation is the multi-agent equivalent of input sanitisation.

From automation idea to production agent.

  1. 01. Agent Scoping & Design

    Task decomposition mapping, tool requirement specification, autonomy boundary definition, human-in-the-loop gate design and failure mode enumeration.

  2. 02. Architecture Design

    Reasoning strategy selection, state machine design, tool interface specification, memory architecture, multi-agent topology (if applicable) and evaluation framework design.

  3. 03. Tool Integration Engineering

    Tool definition development with full input/output schemas, error handling for all failure modes, retry logic, timeout management and safe execution boundaries.

  4. 04. Agent Core Development

    Core agent implementation — task planner, execution loop, memory integration, tool orchestration, error recovery and human escalation path.

  5. 05. Evaluation Framework

    Automated evaluation suite — task completion rate, path efficiency, tool call accuracy, edge case handling, cost tracking and regression test suites.

  6. 06. Safety & Guardrails Review

    Prompt injection testing, tool misuse scenario testing, data leakage assessment, autonomy boundary verification and human-in-the-loop flow validation.

  7. 07. Production Deployment & Monitoring

    Production deployment with job queue infrastructure, real-time monitoring, anomaly alerting, cost dashboards, escalation workflows and performance optimisation.

AI agent systems we have shipped to production.

Three agent implementations handling real business workflows autonomously.

  • Multi-agent research pipeline automating due diligence for a PE fund

    Challenge: Private equity fund spending 40 hours per deal on initial due diligence needed an agent system that could autonomously research a company — company background, financials, competitors, news, management, risks — and produce a structured memo.

    Architecture: LangGraph orchestrated multi-agent pipeline: Researcher agent (web search + document retrieval), Analyst agent (financial ratio extraction and comparison), Risk agent (news sentiment + litigation search), Writer agent (structured memo generation). Human review at final stage.

    Outcome: Due diligence time reduced from 40 hours to 6 hours per deal, 95% analyst satisfaction with memo quality, 3x more deals evaluated per analyst per month.

  • Autonomous customer support agent handling 78% of tier-1 queries without escalation

    Challenge: SaaS company with 50,000 customers needed to scale support without linear headcount growth — an AI agent that could resolve tier-1 queries (password resets, billing questions, feature how-tos, account issues) autonomously.

    Architecture: RAG system over product documentation and past tickets, tool-use for account lookups and Zendesk actions, escalation classifier routing complex queries to human agents, conversation history for multi-turn context.

    Outcome: 78% autonomous resolution rate, CSAT 4.2/5 for agent-handled tickets (vs 4.6 human), £340K annual support cost reduction, average resolution time 2m 40s.

  • AI trading signal agent monitoring 200 assets and generating structured trade ideas

    Challenge: Proprietary trading desk needed an AI agent monitoring 200 crypto and equity assets continuously — scanning for pattern setups, news catalysts, technical signals — and generating structured trade ideas with entry, stop, target, and rationale.

    Architecture: Asset monitoring pipeline pulling OHLCV data every 5 minutes, technical indicator calculation (RSI, MACD, Bollinger), news feed ingestion, LLM-based signal synthesis generating structured trade ideas, Slack delivery with confidence score.

    Outcome: 200 assets monitored 24/7, 15-20 trade ideas per day, analyst review time reduced 80%, 3 winning trades in first month with 2.1:1 average reward-to-risk.

Frequently Asked Questions

What is the difference between an AI agent and a chatbot?

A chatbot responds to queries within a conversation — single-turn or multi-turn, but reactive. An AI agent executes multi-step tasks, uses tools (APIs, databases, browsers, code execution), maintains state across steps, makes decisions about what to do next and can complete tasks taking minutes to hours without human intervention. The defining characteristic is autonomous action, not conversation. Most products described as 'agents' today are sophisticated chatbots with tool calling.

How do you make agents reliable enough for production?

Production agent reliability comes from: strict tool definitions with explicit input/output schemas and error handling, well-defined failure modes with specific recovery behaviours, human-in-the-loop gates for irreversible actions, comprehensive evaluation frameworks that test edge cases and regression suites that run on every deployment. The agent architecture — how it plans, handles tool failures and knows when to escalate — is more important than the model powering it.

What is indirect prompt injection and how do you prevent it?

Indirect prompt injection occurs when malicious instructions are embedded in content the agent retrieves — a webpage, a document, a database record — causing the agent to follow attacker-controlled instructions rather than the intended task. Prevention: treat all tool outputs as untrusted data, not system context; use structured output schemas that separate data from instructions; implement output filtering that detects instruction-like patterns in retrieved content; and limit tool access scope so compromised retrieval cannot trigger high-risk actions.

When should tasks require human approval?

Design a tiered autonomy model based on reversibility and impact: read-only tasks (search, retrieve, analyse) can be fully autonomous. Write actions with bounded scope (create draft, update a record) can be autonomous with logging. Irreversible or high-impact actions (send email, delete data, make payment, post publicly) require human approval. The boundary must be implemented in code with explicit gate definitions — not managed by prompting the model to 'ask before doing important things'.

How do you evaluate whether an agent is working correctly?

Agent evaluation requires: a test suite of representative tasks with defined success criteria and expected tool call sequences; task completion rate measured against the full test suite; path efficiency metrics (did the agent take unnecessary steps?); tool call accuracy (did it call the right tools with correct parameters?); and regression testing after every model update or prompt change. LangSmith, Arize Phoenix and custom evaluation pipelines provide the infrastructure. Without continuous evaluation, you cannot detect quality degradation.

How do you control agent costs at scale?

Agent cost control requires: model tier routing (cheap model for planning, expensive model for reasoning steps requiring deep understanding); context window management (prune irrelevant history, retrieve only necessary documents); tool call caching (cache deterministic tool results); async execution for non-interactive tasks (use batch inference pricing); and per-user or per-task cost budgets with hard limits. A naively implemented agent with a long context window and many tool calls can cost $1-5 per task execution. Well-designed cost controls bring this to $0.05-0.30.

What agent frameworks do you use?

Framework selection follows task complexity and reliability requirements. LangGraph for stateful multi-step agents with complex branching logic. LangChain for RAG-integrated agents with extensive tool libraries. AutoGen or CrewAI for multi-agent systems with role specialisation. Custom orchestration for production systems where framework abstractions create debugging complexity or version instability. We start with the framework that accelerates development and migrate to custom orchestration where production requirements outgrow framework limitations.

Can agents work with our existing APIs and tools?

Yes — any REST API with an OpenAPI specification can be converted to an agent tool definition. We also build custom tool definitions for internal APIs, databases, browser automation and code execution environments. The quality of tool definitions — input/output schemas, error handling, retry logic — directly determines agent reliability. A poorly defined tool is a reliability gap in every task that uses it.

How do you handle agents that get stuck in loops?

Loop detection requires: step count limits with hard cutoffs, repetition detection in tool call sequences, progress tracking that measures whether the agent is moving toward task completion, and watchdog timers that escalate to human review after a defined period of non-progress. ReAct-style agents are particularly prone to reasoning loops on ambiguous tasks. Explicit replanning triggers, 'ask for clarification' tool definitions and task decomposition validation during planning reduce loop frequency.

How long does a production AI agent take to build?

A single-purpose production agent with 3-5 tools, evaluation framework and human-in-the-loop gates: 6-10 weeks. A multi-agent system with specialised roles and complex orchestration: 12-20 weeks. The evaluation framework and safety review typically take 30-40% of the total build time — teams that skip this ship agents they cannot measure or maintain. We deliver in milestone-based sprints with staged capability releases.