7th October 2026 · 15 min
How to Build AI Agents That Actually Work
Only 17% of UK businesses using AI have any AI policy, and just 5% have a formal written policy. That means the most important part of building an AI agent isn't choosing a clever model, it's controlling what the system can access, decide and do.
The popular advice is to define a prompt, connect a model to a few tools and call the result an agent. That approach can produce an impressive demonstration in an afternoon. It can also create a system that invents an answer, repeats an expensive tool call, exposes sensitive information or takes an irreversible action without anyone knowing who approved it.
Production agents need a narrower definition of success. They must solve a real operational problem, use dependable business data, operate within explicit permissions and hand uncertain work to a person. The language model is only one component. Workflow design, observability and governance determine whether the agent earns trust.
Table of Contents
- Why Most AI Agents Fail Before Production
- Building a Controlled Workflow Before Adding Intelligence
- Designing Memory Architecture That Doesn't Leak Context
- Choosing Orchestration Patterns for Different Complexity Levels
- Measuring Real Cost and Value Beyond Token Counts
- From Shadow Mode to Production - The Deployment Checklist
Why Most AI Agents Fail Before Production
The policy gap is a practical warning, not an administrative footnote. The UK Government's AI Adoption Research reports that only 17% of UK businesses using AI had any AI policy, while just 5% had a formal written policy. A team can have a capable model and still lack an owner, an approval process, an incident route or a defensible record of how the agent reached a decision.
That gap appears when a prototype meets real work. A support agent that drafts an answer from a clean test document behaves differently when a customer's account is incomplete, two policies conflict or a request requires access to a billing system. Prompt quality might improve the response, but it won't decide whether the agent should refund an order, reveal account information or escalate the case.

Capability isn't authority
The first design question shouldn't be “Which model is smartest?” It should be “Which decisions may this system make without approval?” That distinction changes the architecture. A capable agent with broad access can create more exposure than a less capable agent restricted to retrieving information and preparing a draft.
Assign a business owner before writing the system prompt. That person should define the intended outcome, approve the source systems, set escalation rules and decide what evidence must be retained. Engineering can implement the controls, but it shouldn't have to guess whether a customer-facing action is acceptable.
Practical rule: Give an agent the minimum authority required for its job, then expand access only when evaluation evidence justifies it.
The Office for National Statistics analysis of UK firms found that 9% of UK firms had adopted AI in 2023, compared with 69% using cloud-based computing. It also found that 88% of firms in the top decile of management-practice scores had adopted at least one advanced technology, compared with 51% in the bottom decile. The useful lesson isn't that management replaces engineering. It's that documented processes and clear ownership make technical adoption more dependable.
A production review should therefore examine more than generated text:
- Ownership: Someone can approve changes and make decisions during an incident.
- Authority: Each tool has a defined purpose, permission scope and approval requirement.
- Evidence: Logs show the input, retrieved sources, tool calls, output and human intervention.
- Measurement: The team tracks operational outcomes, not just whether an answer sounds convincing.
A weekend project usually stops at a successful conversation. A business system must also explain what happened, why it happened and who can stop it.
Building a Controlled Workflow Before Adding Intelligence
Start with one measurable business outcome. “Build an AI employee” is too broad to test, budget or govern. “Classify incoming support requests, identify the relevant policy and prepare a response for review” gives the team a defined input, a bounded process and an observable result.
Before choosing an orchestration library, write down the workflow in ordinary language. Identify the systems the agent needs, the information it may read, the actions it may request and the moments when a person must approve or take over. This problem-first approach is also consistent with practical guidance on AI workflow automation for small businesses, where a focused process is easier to introduce than an unrestricted assistant.
Define the contract
A useful agent contract contains five parts:
- The task: State the single operational job and the conditions under which it applies.
- The inputs: Name the fields, documents, messages or records the agent may use.
- The tools: List each permitted action, its arguments and its failure response.
- The escalation: Specify what happens when information is missing, confidence is low or the request is sensitive.
- The outcome: Choose measures such as resolution time, escalation rate, factual accuracy, cost per task and customer satisfaction.
Don't let the model invent the contract at runtime. Return structured JSON, validate it against a schema and reject malformed output before it reaches a downstream service. A support classification might require a request category, urgency, evidence references and an escalation flag. If one field is absent, the system should stop safely rather than improvise.
Make failure a designed state
Reliable agents behave predictably when dependencies fail. Use timeouts for model and tool calls, bounded retries for transient errors and idempotent operations so a retry doesn't create duplicate tickets or send the same message twice. Record an audit event for every retrieval, decision and action, including the identity of the user or service account that initiated the task.
Retrieval should return source-level citations, not only a block of text. The reviewer needs to see which policy passage supports the draft, and the system needs a way to distinguish an answer grounded in approved material from one produced without evidence.
Evaluate the workflow with a test set containing ordinary requests, ambiguous language, adversarial instructions and personally identifiable information cases. Track task success, grounded-answer rate, unsafe-action rate, tool-call accuracy, latency, cost per completed task and human override rate. Run the agent in shadow mode first. It can make recommendations beside the existing process, but it shouldn't change live records until the team agrees that its behaviour meets release criteria.
This is slower than connecting a model directly to an inbox. It's faster than repairing an uncontrolled system after it has made a customer-facing mistake.
Designing Memory Architecture That Doesn't Leak Context
Memory isn't a feature you switch on. It's a set of retention and retrieval decisions. An agent should receive the context needed for the current task, not a growing archive of everything a customer or employee has ever said.
Separate memory by purpose. Session context holds the short-lived information needed to complete the current interaction. A structured profile stores stable, permissioned attributes such as account status or preferred language. A document index supports retrieval from unstructured material such as policies, manuals and contracts. These stores have different update rules, access controls and retention requirements.

Match storage to the question
A vector store is useful when the agent needs to find semantically related passages in a large document collection. It's a poor substitute for a transactional database when the question is “What is this customer's current balance?” Use structured storage for values that need exact filtering, updates and access checks.
A session cache can help the agent maintain continuity during a conversation. It shouldn't become a permanent profile by accident. If the system stores a user's temporary frustration as a lasting preference, later responses may be biased by information that was never meant to persist.
The same principle applies to retrieval scope. Restrict searches by tenant, user role, document status and task purpose before the model sees the results. A customer-support agent shouldn't retrieve internal investigation notes just because those notes contain terms similar to the customer's question.
More context doesn't automatically create a better answer. It often creates a harder access-control problem.
The UK Business Data Survey 2026 reported AI use among 41% of surveyed businesses handling digitised data, while governance maturity remained low. That combination makes memory boundaries especially important. Organisations are connecting agents to more information before they have fully decided what should be retained, who may retrieve it or how an old record should be removed.
Add traceability and expiry
Every retrieved fact should carry its origin, document version and relevant access context. Store citations with the agent's intermediate result so a reviewer can trace a response back to an approved source rather than searching through a collection manually.
Define retention rules before implementation:
- Session expiry: Decide when conversational context is deleted or made inaccessible.
- Profile updates: Require an explicit event or approved workflow before changing durable attributes.
- Personal data: Minimise collection, restrict access and test requests involving personally identifiable information.
- Document changes: Re-index superseded policies carefully and preserve the version used for earlier decisions.
These choices belong in the application architecture, alongside queues, APIs and databases. Guidance on web application architecture is relevant here because an agent is still software with identity, storage, dependencies and failure modes.
A useful memory design makes the agent less dependent on vague recollection. It retrieves the smallest relevant evidence set, verifies the structure of that evidence and discards context when the task ends.
Choosing Orchestration Patterns for Different Complexity Levels
An agent doesn't become more useful because its architecture contains more agents. Complexity introduces more prompts, more state transitions, more failure points and more cost. Choose the simplest orchestration pattern that can express the workflow and make its decisions observable.
Linear chains
A linear chain suits a predictable sequence. The system might classify a request, retrieve approved guidance, draft a response and produce a review task. Each stage passes a defined object to the next, which makes schema validation and debugging straightforward.
Use this pattern when the route is stable and the number of decisions is limited. It breaks down when the agent needs to choose among many specialists or revisit an earlier stage. Adding ad hoc branches to a linear chain often produces a difficult-to-test tangle.
Hierarchical routing
A router is appropriate when the first decision determines which specialist workflow should run. An incoming support message might go to billing, technical troubleshooting or account access. Each specialist can have narrower tools and source permissions than a general-purpose agent.
This pattern gives the team a useful control point. The router can reject unsupported categories, apply an initial risk rule and send sensitive cases directly to a human. Its weakness is routing error. If the first classification is wrong, the specialist may produce a polished answer to the wrong problem, so the routing result needs its own evaluation.
Multi-agent systems
Multiple agents make sense when the task benefits from parallel work, negotiation or independent review. For example, separate workers might research several approved sources before a reviewer reconciles their findings. They don't make sense when one constrained workflow could complete the job with fewer moving parts.
UK adoption data shows why this restraint matters. The AI Adoption Research report reports AI use among 36% of large businesses, 23% of medium-sized businesses and 14% of microbusinesses. A smaller organisation usually gains more from a focused agent for triage, document summarisation or internal task routing than from a broad autonomous system that requires extensive integration and supervision.
| Pattern | Use it when | Main control concern |
|---|---|---|
| Linear chain | The task follows a stable sequence | Validate every hand-off |
| Hierarchical router | Different cases require different specialists | Test misrouting and fallback |
| Multi-agent system | Parallel reasoning or independent review adds real value | Control shared state and compounded cost |
Start with the narrowest viable pattern. Measure reliability and financial value in the target workflow before introducing parallel agents or wider autonomy. A system that completes one repeatable process consistently is more valuable than an elaborate architecture that nobody can explain during an incident.
Measuring Real Cost and Value Beyond Token Counts
Token usage is only the visible part of an agent's bill. The operational total also includes tool calls, retries, browsing or API charges, infrastructure, monitoring and the time a person spends reviewing or correcting the result. A cheap model can become expensive when it loops, retrieves too much context or sends uncertain work to a skilled employee.
Recent UK evidence describes this visibility problem. KPMG's UK AI adoption reporting says 30% of UK leaders report difficulty with usage-based AI costs, 42% have only partial visibility into AI spending and 33% cite limited understanding of token-related cost structures as a deployment challenge.
Budget the complete task
Set a cost ceiling for each workflow execution. The ceiling should cover the model budget, retrieval, tool calls, retries and expected human review. If the agent reaches that limit, it should stop, record the reason and escalate rather than continue searching indefinitely.
Track costs by completed outcome, not by activity. Ten tool calls that produce a correct resolution are not equivalent to ten tool calls that end in manual rework. Useful operational fields include:
- Execution cost: Model, retrieval, API and infrastructure charges for one run.
- Review cost: Time spent checking, editing or replacing the output.
- Failure cost: Reversed actions, duplicate work, customer impact and incident handling.
- Outcome value: The business result that the workflow was designed to produce.
Compare the agent with a deterministic workflow or conventional software baseline. If a rules-based classifier handles the common cases and routes exceptions cleanly, it may be more economical than an agent that reasons over every request. An agent earns its place when its flexibility creates enough value to justify its variable cost and control burden.
Budgeting principle: Automation rate is not the same as profitable automation.
Know when not to use an agent
An agent may be economically inappropriate when exceptions are frequent, data retrieval is expensive or errors require substantial manual correction. It may also be the wrong choice for a stable process that a form, validation rule or scheduled integration can handle more reliably.
A narrow first workflow creates a clearer procurement decision. The team can define the expected volume, maximum cost per task, review requirement and acceptable failure modes before it expands the system. It can also stop an unproductive experiment without having to unwind a general-purpose “coworker” embedded across the business.
Value measurement should remain tied to the original outcome. If the goal is faster support resolution, monitor resolution quality and escalation workload. If the goal is document processing, measure accurate completed records and correction time. A successful agent is one that improves the process at an acceptable risk and cost, not one that produces the longest answer.
From Shadow Mode to Production - The Deployment Checklist
Production starts after the demo has stopped touching live systems. In shadow mode, the agent receives realistic work beside the existing process, generates its proposed classification or action and leaves the current workflow unchanged. Reviewers compare both paths without allowing the agent to create an irreversible customer or financial event.
Define the release thresholds before looking at the results. The team should agree how much factual accuracy is acceptable, what level of unsafe-action rate blocks release, how much latency the user experience can tolerate and when a human override is mandatory. The exact thresholds depend on the workflow, but the decision must be explicit.
Release in permission stages
A sensible rollout increases authority gradually:
- Read-only evaluation: The agent retrieves approved information and produces recommendations.
- Human-approved actions: A reviewer confirms each permitted tool call.
- Restricted automation: The agent acts only on low-risk cases that pass validation.
- Monitored operation: Alerts, audit records and an immediate escalation route remain active.
Keep an incident owner on the rota. Document rollback procedures, access revocation, data-handling rules and the evidence required to investigate a disputed decision. The UK Government's AI Adoption Research identifies limited AI skills as a major barrier, affecting 60% of businesses in one measure. That makes runbooks and ownership essential, particularly for SMEs without a large platform team.
A realistic readiness decision
Consider a support agent that classifies incoming messages, retrieves an approved policy and drafts a response. During shadow mode, the team reviews task success, grounded-answer rate, unsafe-action rate, tool-call accuracy, latency, cost per completed task and human override rate. It also checks whether the agent escalates sensitive requests instead of trying to resolve them.
The agent doesn't move to customer-facing work because its responses sound natural. It moves when the agreed quality and safety thresholds are met across normal, ambiguous, adversarial and personal-data test cases. The team starts with read-only access, enables a human approval step for outbound messages and keeps automatic escalation for cases that lack supporting evidence.
Monitoring continues after launch. Alert on unusual retry volume, unexpected tool usage, rising overrides, missing citations and cost anomalies. Review a sample of completed tasks, update the test set with real failures and keep a rollback path that a named person can activate without waiting for a new model release.
The build process should also cover the surrounding product, permissions and operational interface. Teams evaluating web application development services need to treat the agent as part of the application rather than a detached chat window.
The difference between a weekend project and a production agent is visible in the controls. It has a defined owner, bounded authority, traceable evidence, measured economics and a safe way to stop.
Digital Souls Studios LTD builds web applications, SaaS products, AI automation and custom support or operations agents, taking projects from specification through launch and ongoing support. If you want to turn a repeatable UK business workflow into a governed agent with measurable controls, visit Digital Souls Studios LTD to discuss the process.