An agent that works 90% of the time is not a product. It is a liability.
Agent reliability is architectural. Task decomposition strategy, tool call reliability, context window management, error handling and recovery, human-in-the-loop design, evaluation at scale — each a separate engineering problem. Most agent demos work in the happy path. Production agents need edge case handling, graceful retry logic, deterministic escalation to humans, and evaluation frameworks that catch regressions before users do. PROPELOO builds agents designed for the 10% of cases that break everything.
Frequently Asked Questions
What is the difference between an AI agent and a chatbot?
A chatbot responds to queries within a conversation — single-turn or multi-turn, but reactive. An AI agent executes multi-step tasks, uses tools (APIs, databases, browsers, code execution), maintains state across steps, makes decisions about what to do next and can complete tasks taking minutes to hours without human intervention. The defining characteristic is autonomous action, not conversation. Most products described as 'agents' today are sophisticated chatbots with tool calling.
How do you make agents reliable enough for production?
Production agent reliability comes from: strict tool definitions with explicit input/output schemas and error handling, well-defined failure modes with specific recovery behaviours, human-in-the-loop gates for irreversible actions, comprehensive evaluation frameworks that test edge cases and regression suites that run on every deployment. The agent architecture — how it plans, handles tool failures and knows when to escalate — is more important than the model powering it.
What is indirect prompt injection and how do you prevent it?
Indirect prompt injection occurs when malicious instructions are embedded in content the agent retrieves — a webpage, a document, a database record — causing the agent to follow attacker-controlled instructions rather than the intended task. Prevention: treat all tool outputs as untrusted data, not system context; use structured output schemas that separate data from instructions; implement output filtering that detects instruction-like patterns in retrieved content; and limit tool access scope so compromised retrieval cannot trigger high-risk actions.
When should tasks require human approval?
Design a tiered autonomy model based on reversibility and impact: read-only tasks (search, retrieve, analyse) can be fully autonomous. Write actions with bounded scope (create draft, update a record) can be autonomous with logging. Irreversible or high-impact actions (send email, delete data, make payment, post publicly) require human approval. The boundary must be implemented in code with explicit gate definitions — not managed by prompting the model to 'ask before doing important things'.
How do you evaluate whether an agent is working correctly?
Agent evaluation requires: a test suite of representative tasks with defined success criteria and expected tool call sequences; task completion rate measured against the full test suite; path efficiency metrics (did the agent take unnecessary steps?); tool call accuracy (did it call the right tools with correct parameters?); and regression testing after every model update or prompt change. LangSmith, Arize Phoenix and custom evaluation pipelines provide the infrastructure. Without continuous evaluation, you cannot detect quality degradation.
How do you control agent costs at scale?
Agent cost control requires: model tier routing (cheap model for planning, expensive model for reasoning steps requiring deep understanding); context window management (prune irrelevant history, retrieve only necessary documents); tool call caching (cache deterministic tool results); async execution for non-interactive tasks (use batch inference pricing); and per-user or per-task cost budgets with hard limits. A naively implemented agent with a long context window and many tool calls can cost $1-5 per task execution. Well-designed cost controls bring this to $0.05-0.30.
What agent frameworks do you use?
Framework selection follows task complexity and reliability requirements. LangGraph for stateful multi-step agents with complex branching logic. LangChain for RAG-integrated agents with extensive tool libraries. AutoGen or CrewAI for multi-agent systems with role specialisation. Custom orchestration for production systems where framework abstractions create debugging complexity or version instability. We start with the framework that accelerates development and migrate to custom orchestration where production requirements outgrow framework limitations.
Can agents work with our existing APIs and tools?
Yes — any REST API with an OpenAPI specification can be converted to an agent tool definition. We also build custom tool definitions for internal APIs, databases, browser automation and code execution environments. The quality of tool definitions — input/output schemas, error handling, retry logic — directly determines agent reliability. A poorly defined tool is a reliability gap in every task that uses it.
How do you handle agents that get stuck in loops?
Loop detection requires: step count limits with hard cutoffs, repetition detection in tool call sequences, progress tracking that measures whether the agent is moving toward task completion, and watchdog timers that escalate to human review after a defined period of non-progress. ReAct-style agents are particularly prone to reasoning loops on ambiguous tasks. Explicit replanning triggers, 'ask for clarification' tool definitions and task decomposition validation during planning reduce loop frequency.
How long does a production AI agent take to build?
A single-purpose production agent with 3-5 tools, evaluation framework and human-in-the-loop gates: 6-10 weeks. A multi-agent system with specialised roles and complex orchestration: 12-20 weeks. The evaluation framework and safety review typically take 30-40% of the total build time — teams that skip this ship agents they cannot measure or maintain. We deliver in milestone-based sprints with staged capability releases.