AI
Building Production-Ready Agentic AI
- Author
- Deploint Engineering
- Published
- Reading time
- 7 min read
- Sections
- 08
An agent that books a meeting or files a ticket in a demo takes a few dozen lines of code. An agent that does the same thing reliably, for thousands of users, against real systems with real permissions, is an engineering program. The gap between the two is rarely model quality. It is everything around the model: how tools are defined, how state is managed, how failures are contained and how behavior is measured over time.
This article sets out the architecture decisions we consider essential before an agentic system handles production work, and the failure modes that tend to appear when they are skipped.
01Decide whether you need an agent
Agentic designs give a model control over the sequence of steps. That flexibility is useful when the path through a task genuinely varies: investigating an incident, reconciling records with inconsistent formats, or assembling an answer from several systems. It is a liability when the path is known. If a process always runs the same five steps, a deterministic workflow that calls a model for specific judgments will be cheaper, faster, easier to test and easier to explain.
A useful test is to draw the task as a flowchart. If the flowchart is short and stable, build a workflow. If it branches heavily on information that is only discovered during execution, an agent may earn its complexity. Many production systems end up as hybrids: a fixed outer workflow with a bounded agentic step inside it.
02Treat tools as the real interface
In an agentic system, tools are where the model touches the world. Their design determines most of the system's reliability and almost all of its risk.
- Keep each tool narrow. A tool called
update_shipping_addresswith three typed parameters is safer, and easier for a model to use correctly, than a genericrun_sqltool. - Write tool descriptions as carefully as an API reference for a new engineer: what the tool does, when not to use it and what each error means.
- Validate every argument on the server. Model output is untrusted input, however well it usually behaves.
- Return compact, structured results. Large raw payloads consume context and make later steps less reliable.
- Make write operations idempotent and attach a request identifier, so a retried step cannot create a duplicate order or a second refund.
Tools should also carry the identity of the user the agent is acting for. An agent that runs under a shared service account can reach data the requesting user could not, which turns every prompt injection into a potential privilege escalation. Delegated, scoped credentials keep the agent's authority equal to or narrower than the person's.
03Make state explicit
Many agent frameworks keep state implicitly in the conversation history. That works for short tasks and breaks for long ones. Context windows fill, early instructions lose influence and the system cannot be resumed after a failure.
Production agents benefit from an explicit task state held outside the model: the goal, the plan, completed steps, pending approvals and intermediate results. At each step the model reads a compact view of that state and proposes the next action; the orchestrator executes it and records the result. This gives you checkpoints to resume from, a natural place for human approval and a record that can be inspected afterwards.
Bound every loop
Agents can loop: retrying a failing tool, re-planning the same step or oscillating between two options. Set hard limits on steps, tool calls, elapsed time and spend per task, and define what happens when a limit is reached. Usually that means handing the task to a person with the state and trace attached, rather than failing silently.
04Place human review deliberately
Human review is most valuable at points of irreversibility: sending an external message, moving money, changing a record of authority or deleting anything. Requiring approval for every step trains reviewers to click through; requiring it for none moves the risk onto customers. Classify each tool by consequence and reversibility, and attach approval requirements to the tool, not to the prompt.
05Contain prompt injection structurally
Any content an agent reads, including emails, web pages, documents and tool results, can contain instructions. Filters that try to detect malicious text help, but they cannot carry the load alone. Structural controls matter more:
- Limit what the agent can do after reading untrusted content. An agent that has just processed an inbound email should not hold a tool that sends data outside the organization in the same task.
- Separate data from instructions in prompts, and label untrusted content clearly.
- Allow-list destinations for outbound actions: domains, recipients and accounts.
- Log the full chain of inputs behind each action, so an injected instruction can be traced.
06Evaluate continuously, not once
Agent behavior depends on the model, the prompts, the tool descriptions and the data tools return. A change to any of them can change outcomes, so evaluation belongs in the delivery pipeline.
- Build a task suite from real cases: inputs, expected end states and actions that must never happen. Include the awkward cases, not only the common ones.
- Score outcomes, not wording. Check whether the right record changed in the right way, whether approval was requested where required and whether the agent stopped when it should have.
- Run the suite on every change to prompts, tools, models or retrieval configuration, and block releases on regressions.
- Sample production traces for human review, and add confirmed failures to the suite.
Model-graded evaluation helps at scale, but calibrate it against human judgment on a sample, and re-check that calibration whenever the grading model changes.
07Instrument every step
When an agent does something unexpected, the first question is why. Answering it requires a trace of every step: the state the model saw, the action it proposed, the tool call and arguments, the result, the elapsed time and the cost. OpenTelemetry-style tracing, with a span per step and per tool call, lets agent traces sit alongside the rest of your service telemetry. Redact sensitive fields at the point of capture, not in the dashboard.
Useful operational signals include task completion rate, escalation rate, steps per task, tool error rates, approval rejection rate and cost per completed task. Movement in these often reveals a problem before users report it.
08A production readiness checklist
Checklist10 checks
Before an agent handles production work
- 01The task needs dynamic planning; a fixed workflow was considered and rejected for stated reasons.
- 02Every tool is narrow, typed, validated on the server and idempotent for writes.
- 03The agent acts with the requesting user's delegated permissions, not a shared privileged account.
- 04Task state lives outside the model, with checkpoints and resumability.
- 05Step, time and spend limits are enforced, with a defined handoff when they are reached.
- 06Irreversible actions require human approval, with the evidence shown.
- 07Untrusted content cannot directly trigger outbound or destructive actions.
- 08An evaluation suite runs on every change and gates release.
- 09Every step is traced, with sensitive data redacted at capture.
- 10There is a kill switch, and someone on call knows how to use it.
None of these controls is novel. They are the disciplines that make any distributed system dependable: narrow interfaces, explicit state, bounded retries, least privilege, testing and observability. Agentic AI does not replace that engineering. It raises the cost of skipping it.