Guide

AI agents for business: where they fit, and the controls they need

An AI agent is software that uses an AI model to decide its own next steps and to call tools, such as a search, a database or a business application, until a task is done. Agentic AI is the wider label for systems built this way. A multi-agent system splits one job across several agents, and orchestration is whatever decides which agent runs next and when the job stops. In a business, agents pay back on work that needs judgement across messy inputs and has resisted rules-based automation, such as triaging requests or matching documents that arrive in different formats. They are the wrong tool where a fixed sequence of steps would do. Before an agent touches a real system it needs its own identity with limited access, automatic checks on every action it proposes, a person to approve anything consequential, limits on steps and spend, and a record of what it did.

AI agent. Software that uses an AI model to choose its own next steps and call tools, such as searches, databases and business applications, to complete a task on someone’s behalf.

1AYM is an OpenAI Select Partner. Its founder holds personal Claude certifications. 1AYM is not an Anthropic partner.

Checked . Definitions and vendor product descriptions were read on the publishers’ own pages. Vendors change their agent products and documentation without notice, so confirm the details on the sources below before you rely on them.

AI agents, agentic AI, multi-agent systems and orchestration

The four terms are used loosely, vendors included. These are the working definitions this guide uses, each tied to a published source listed at the end of the page.

AI agent
Software that uses an AI model to decide its next step and to call tools until a task is done. OpenAI’s guide to building agents describes them as systems that complete tasks on a user’s behalf with a high degree of independence, and says that applications which use a model without letting it control the workflow, such as a simple chatbot or a sentiment classifier, are not agents.
Agentic AI
An umbrella term, not a product. Anthropic calls everything built this way an agentic system, and separates workflows, where code fixes the path the model follows, from agents, where the model directs its own process and tool use. The UK Information Commissioner’s Office describes agentic AI as generative AI combined with additional tools and new ways of interacting with the world, which lets it automate more open-ended tasks.
Agentic AI vs AI agents
An AI agent is one working unit. Agentic describes how much a system decides for itself, so a fixed workflow with one model step and a free-running agent are both agentic, to different degrees. For a buyer the useful question is not which label a vendor uses but how many of the steps the model chooses, because that decides both the cost and what has to be checked.
Multi-agent system
Several agents, each with its own instructions and tools, working on one job. Anthropic’s research system uses a lead agent that plans the work and hands parts of it to subagents running in parallel. OpenAI’s guide describes two shapes: a manager agent that calls the others as tools, and peers that hand the work from one to the next.
Agent orchestration
Whatever decides which agent runs, in what order, and when the job is finished. The OpenAI Agents SDK documentation separates orchestration by the model from orchestration in code, and says code makes speed, cost and performance more predictable. Microsoft’s Agent Framework ships sequential, concurrent, handoff, group chat and manager-led patterns.

Where AI agents work in a business

Agents earn their cost on work that needs judgement across inputs no rule set captures, and where a wrong step can be caught before it lands. OpenAI’s guide names three signs: decisions that need judgement and are full of exceptions, rule sets that have grown too costly to maintain, and heavy reliance on unstructured material such as documents and conversations. Where a use case does not clearly meet those tests, the same guide says a deterministic solution may be enough.

The table below is our reading of how that plays out across common business functions. It describes where to look, not results anyone has measured in your business.

Where an agent tends to fit in each business function, and where a simpler tool usually does the job
DimensionWorth an agentUsually better without one
Service and operationsTriaging incoming requests, drafting the reply and gathering the records a person needs to decideRouting on one field, such as a form’s category, which a rule does exactly
FinanceMatching invoices, statements and remittances that arrive in different formats, with every proposed entry checked against the ledgerPosting entries from a structured feed, which a scheduled job does the same way every time
Data and reportingAnswering a question in plain words against an agreed set of business definitionsA monthly report with fixed figures, which a dashboard already produces
Compliance and riskReading contracts or policies and flagging the clauses a reviewer should look atA decision that must be explained rule by rule, where a rules engine is clearer
Software engineeringCoding agents working in a repository, with every change going through the normal review pathChanges to production data or to access, made without review

One question comes before all of them: does the model need to choose the steps at all? Anthropic’s published advice is to find the simplest solution possible and add complexity only when it is needed, and it notes that agentic systems often trade speed and cost for better task performance. Many useful systems are a fixed workflow with one model step inside it, which is cheaper to run and easier to test than an agent.

Where agents are the wrong tool

Four situations argue against an agent. When the steps are known and stable, a workflow or a scheduled job does the same work more cheaply and more predictably. When nobody can say what a correct result looks like, nothing can check the agent either. When every output is read by a person before anything happens, a drafting assistant is enough, because the reader is the check. And when an error would land before anyone could see it, the agent should not be allowed to act until the controls in the next section exist.

Multi-agent systems add a further cost. Anthropic reports that, in its data, agents use about four times the tokens of a chat interaction and multi-agent systems about fifteen times, and that work where every agent needs the same context, or where the agents depend heavily on one another, is not a good fit for multi-agent systems today. OpenAI’s guide likewise recommends getting as far as possible with a single agent before adding more.

The controls an agent needs before it touches a real system

Accuracy is the wrong measure for an agent. What matters is what happens on the occasions it is wrong, and whether anything notices before the consequence lands. The controls below are the ones to have in place before an agent may write to anything that matters. The agent proposes and deterministic checks decide what proceeds; anything that fails a check goes to a person with the reason attached. The mechanism is set out in full in our published pattern for gating an agent’s actions.

Controls to set before an agent can act on a real system, what each one stops, and the record it leaves
DimensionWhat it stopsEvidence it leaves
Its own identity, with least accessThe agent reaching data or systems the person it acts for could notIdentity configuration and access review results
Checks on every proposed actionA plausible but wrong entry being saved: schema, business rules and reconciliation against the system of record run firstA log of each proposed action and the checks it passed or failed
A person for consequential actionsPayments, changes to someone’s access or anything irreversible going ahead unseenThe approver, the time and the reason attached to each held action
Step, time and spend limitsA run that retries without end or runs up a billRun records showing which stop condition ended each run
Tests before every changeA new prompt, model or tool quietly changing behaviourTest results kept for each release
A dry run and a way backA bad first run that cannot be undoneDry-run output reviewed before going live, and a tested rollback
A switch to turn it offAn agent that keeps acting while a problem is investigatedTests that show the switch works

Vendors now ship parts of this in their agent products. The OpenAI Agents SDK includes guardrails that validate an agent’s inputs and outputs, and built-in tracing of each run. Anthropic’s Claude Agent SDK lets a developer set which tools run automatically and which need approval, and run custom code at set points in the agent’s lifecycle. Microsoft’s Agent Framework supports tools that pause a workflow for human review before they run. These are building blocks. Which actions need a check, and who approves them, is still the business’s decision. OpenAI’s own guide names two triggers for handing control to a person: an agent exceeding its failure limits, and actions that are sensitive, irreversible or high-stakes.

Two published engagement files describe parts of this. At a large international marketing agency, we build managed-agent workflows for finance-critical tasks, so deterministic checks decide what proceeds (engagement file D-01). In engagement file D-06, a connector answers questions from finance data while carrying each person’s existing permissions, so nobody sees more than they already could.

Where an agent handles personal data, data protection law applies to it as to any other processing. The ICO’s January 2026 report on agentic AI says organisations remain responsible for the data protection compliance of the agentic AI they develop, deploy or integrate, and the OWASP GenAI Security Project published a Top 10 for Agentic Applications in December 2025 setting out the main security risks. This is not legal advice. Which rules apply to a particular agent is a question for your legal adviser or data protection officer. For where these controls sit in a wider policy, our governance guide shows an AI policy mapped to its controls, clause by clause.

What drives the cost and the risk

No two agent projects cost the same, and a single published price would hide the drivers below. Each one can be estimated before a build starts.

The runtime is increasingly a vendor product rather than something to build. The vendor runtimes we call an AI harness already include an agent loop, tools and permission settings: Anthropic’s Claude Agent SDK, for example, packages the loop, tools and permissions that run Claude Code. What a business still owns is the work around the runtime: the definitions, the checks, the access model and the person who answers when it fails.

How many steps the model chooses
Every step is a model call. A fixed workflow with one model step costs a known amount per run; an agent that plans its own path costs whatever the path turns out to be, which is why stop conditions matter. Anthropic’s figures put an agent at about four times the tokens of a chat, and a multi-agent system at about fifteen.
Volume and context
Cost per run multiplied by runs per day. A design that is cheap at ten users can become expensive at a thousand if each run carries more context than the task needs.
Agreeing what correct means
Connecting an agent to data is often quick. Agreeing what a correct answer is takes longer. On engagement file D-06, connecting the data took an afternoon, and the week that followed, agreeing definitions with the finance director, was the actual work.
The checks and the review queue
Deterministic checks are cheap to run and take engineering time to write. Human review costs staff time for every held item, so the design aims to hold only what fails a check.
What the agent can change
Access sets the risk. Read-only access to public documents needs little; write access to money, identity or customer records needs every control in the table above.
Running it after launch
Monitoring, retesting when a vendor changes a model, and a named owner. An agent nobody owns after launch is a liability, whatever it cost to build.

A checklist for your first agent project

Use this before you commission a build, from us or from anyone else. Each item should have a written answer before any code is written.

1. One workflow, with an owner
Pick one workflow with a named owner, a known volume and a cost you can state today, such as hours a week or the time a request waits.
2. A test for correct
Collect real past cases with their right answers, so the agent can be tested against them before anyone trusts it.
3. The simplest design first
Check whether a fixed workflow with one model step would do. Use an agent only where the steps genuinely change from case to case.
4. Read before write
Start read-only, or with the agent drafting and a person acting, and widen its access only once the results hold up.
5. Controls before go-live
Identity, checks, approvals, limits, tests, a dry run and an off switch, as in the table above, each with a named owner.
6. A cost ceiling
Set a budget per run and per month, and a stop condition that ends any run that goes past it.
7. Data review
Record what data the agent reads and where it is processed, and complete a data protection impact assessment where the use is likely to be high risk. This is not legal advice.
8. A measure and a date
Decide the before-and-after measure and the date you will read it. If the number does not move, change the design or stop.

If the checklist shows the workflow is not ready, that is a useful answer, and cheaper to learn now than after a build. If you want help with it, a fixed-scope AI strategy sprint of typically two to four weeks scores candidate use cases on value, feasibility and risk.

For engineers: agent architecture, controls and evidence

The controls table in engineering terms. Each item can be checked by reading configuration, code or logs.

Orchestration in code first
Prefer a coded workflow with bounded model steps. Where the model chooses the path, cap iterations, tool calls and tokens per run, and fail closed on a breach.
Tool contracts
Give each tool a typed schema and validate its arguments server-side. Make writes idempotent, keyed on a request ID, so a retry cannot double-post.
Identity
Run each agent as its own service principal. Where the source system supports delegation, carry the end user’s identity on tool calls and inherit their entitlements.
Verifier gates
Every proposed write passes schema, business-rule and reconciliation checks against the system of record. A failure holds the action and routes it to a named approver with the reason.
Multi-agent boundaries
Give each subagent the narrowest tool set its job needs. Pass structured results between agents rather than free text, so a gate can check them.
Untrusted input
Treat anything an agent reads from documents, email or the web as data, never as instructions. Gate any tool call whose arguments came from that content.
Evaluation in CI
Keep prompts, model versions, tool definitions and rules in the repository. Each change runs a regression set of real cases against recorded thresholds.
Tracing and audit
Record each run: input references, model and prompt version, tool calls, check results, approver and timestamps. Store it append-only, with a stated retention period.
Switches and dry runs
Put each agent behind a flag that turns it off without a deploy, and give every write path a dry-run mode that shows the diff before it lands.

Sources

  1. [1]Anthropic, Building effective agents (19 December 2024)
  2. [2]Anthropic, How we built our multi-agent research system (13 June 2025)
  3. [3]OpenAI, A practical guide to building agents (PDF, April 2025)
  4. [4]OpenAI Agents SDK documentation: overview
  5. [5]OpenAI Agents SDK documentation: agent orchestration
  6. [6]Anthropic, Claude Agent SDK overview
  7. [7]Microsoft Learn, Workflow orchestrations in Agent Framework (updated 25 August 2026)
  8. [8]ICO, Tech Futures: Agentic AI (8 January 2026)
  9. [9]OWASP GenAI Security Project, Top 10 for Agentic Applications for 2026 (9 December 2025)

Frequently asked questions

What is the difference between agentic AI and AI agents?

An AI agent is a piece of software that uses a model to choose its own steps and call tools. Agentic AI is the broader label for systems that act this way, from a fixed workflow with one model step to a set of agents coordinating with each other. The labels matter less than one question: how many steps does the model choose? That decides the running cost and how much has to be checked.

Does a business need a multi-agent system?

Rarely at the start. OpenAI’s guide recommends getting as far as possible with one agent, Anthropic advises the simplest solution that does the job, and Anthropic reports that multi-agent systems use about fifteen times the tokens of a chat. They make sense for work that splits into independent parallel parts, such as broad research. They fit poorly where every agent needs the same context or the parts depend closely on one another.

What is agent orchestration, and who should own it?

Orchestration decides which agent runs, in what order and when the job stops. It can be left to the model or written in code, and code is more predictable in speed, cost and performance. The team that owns the business process should own the orchestration rules, because they decide which steps may run unattended and which must stop for a person.

Is it safe to connect an AI agent to finance or customer systems?

Only with the controls in place first: its own identity with no more access than the person it acts for, deterministic checks on every write, a named approver for consequential actions, limits on steps and spend, and a record of every run. Start read-only and widen access once the results hold up. The agent should never be the last thing to touch a decision that moves money or changes someone’s access.

How much does an AI agent cost to run?

It depends on how many model calls each run makes, how much context each call carries, how many runs a day, and how many items stop for human review. None of those is fixed by the vendor, so an honest estimate starts from one named workflow and its volume. Set a budget per run and a stop condition before the first live run, so the cost cannot run away while you learn.

Should we build our own agent or use a vendor’s?

Use the vendor’s runtime where one fits, because OpenAI, Anthropic and Microsoft each maintain agent tooling with permissions, tracing or approval steps built in. What a vendor cannot supply is your definitions of a correct answer, your access model and the checks on your systems of record. That part is always yours to build or commission, whichever runtime you choose.

Further

We build these systems for a living. See the engagement files for what that looks like in practice, or write to us if yours is the next one.

Last reviewed · 1AYM