← Blog

AI Agents vs AI Workflows: How to Choose for Production

AI agents and AI workflows solve different problems. A practical guide to the difference, when each one fits, and the hybrid pattern that holds up in production.

"Should this be an agent?" is one of the first questions teams ask when they start building with large language models. It is a good question, but it is usually framed as a choice between two products. It is really a choice about who decides what happens next: your code, or the model.

Getting that choice right early saves a lot of rework. An agent where a workflow would do is harder to test, harder to budget and harder to explain to an auditor. A rigid workflow where the task is genuinely open-ended ends up as a tangle of special cases. This article lays out the difference, when each approach fits, and the pattern we reach for most often in production AI engineering.

What we mean by an AI workflow

An AI workflow is a process with a fixed shape where some steps use a model. The sequence is defined in code: a document arrives, OCR runs, a model extracts fields, business rules validate them, a person approves exceptions, and the result is written to the system of record.

Here is that invoice process as a workflow. Every arrow is decided in code; the model only does the work inside the extraction step.

flowchart LR
  A[Invoice] --> B[OCR] --> C[AI: extract]
  C --> D{Rules:<br/>matches PO?}
  D -- yes --> F[Accounting]
  D -- no --> E[Person reviews] --> F

The model does real work, such as reading, classifying, extracting or drafting, but it does not choose the path. The path is yours. That makes an AI workflow predictable in the ways operations teams care about: you know which steps run, in what order, with which permissions, and you can replay any run.

You can see this shape in the example workflows on this site: document processing, support routing, invoice automation and lead qualification all follow a known sequence with AI inside individual steps.

What we mean by an AI agent

An AI agent is a loop in which the model decides the next action. It is given a goal, a set of tools (search, APIs, database queries, code execution) and some context. It picks a tool, looks at the result, and decides again, until it judges the goal met or hits a limit.

flowchart TD
  G[Goal and context] --> M{Model decides<br/>the next action}
  M -- call a tool --> T[Search, API, database]
  T -- result --> M
  M -- goal met --> R[Answer]
  M -- limit reached --> S[Stop and hand off]

The defining property is that control flow is chosen at runtime by the model. That is what makes agents useful for open-ended tasks, and it is also what makes them harder to operate: the same input can take different paths, the number of steps is not fixed, and failures can compound across steps.

The difference that matters: who owns control flow

Most of the practical trade-offs follow from that one distinction.

AI workflow AI agent
Who picks the next step Your code The model
Path through the task Fixed and known Varies per run
Testing Test each step and the whole path Test outcomes across many possible paths
Cost and latency Predictable per run Varies with the number of steps taken
Typical failure A step returns a bad value The loop wanders, repeats or stops early
Auditability Straightforward: fixed steps, fixed permissions Needs full tracing of every decision and tool call
Best fit Repeatable business processes Open-ended tasks with many possible routes

Neither column is better. They suit different problems, and many real systems need both.

When an AI workflow is the right choice

Choose a workflow when the process is already understood and repeatable, even if individual steps are messy. Signals that point this way:

  • The steps are known. You could draw the process on a whiteboard today, and it looks roughly the same every time.
  • Outcomes need to be consistent. The same invoice should produce the same booking, and the same ticket should reach the same queue.
  • There are hard rules. Approval limits, compliance checks, data residency or segregation of duties are easier to enforce as deterministic code around the model than as instructions to it.
  • You need to explain every result. Auditors, finance teams and regulators ask "why did this happen?", and a fixed path gives a clear answer.
  • Volume is high. Predictable cost per run matters when a process runs thousands of times a day.

Most back-office automation lands here. The hard part is rarely choosing the next step; it is reading unstructured input reliably and connecting to the systems around it, which is integration work more than agent design.

When an AI agent earns its place

Choose an agent when the route to the answer genuinely depends on what the model finds along the way. Signals that point this way:

  • The task is open-ended. Research, investigation and troubleshooting, where the next question depends on the last answer.
  • There are many tools and no fixed order. The right sequence of lookups differs from case to case, and encoding every branch would be brittle.
  • Some variation in the path is acceptable. The result can be checked, or a person reviews it before anything irreversible happens.
  • The value of flexibility beats the cost of unpredictability. A slower, costlier run is fine if it handles cases a fixed workflow never could.

Even then, an agent does not have to be unbounded. The agents that hold up in production are usually narrow: a small set of tools, a clear goal, a step limit and a defined way to stop.

The pattern that holds up: a workflow with agent steps

In practice the most reliable design is often a hybrid: a deterministic workflow on the outside, with bounded agent behaviour inside specific steps. The workflow owns the overall process, permissions and hand-offs. An agent is used only where a step genuinely needs to explore.

A support process shows the idea. The workflow receives the email, classifies it, and routes it by rules you control. Inside the "draft a reply" step, an agent may look up the order, check the help centre and read recent tickets before drafting. The workflow then applies policy checks and sends the draft to a person for approval. The agent's freedom is confined to one step, and nothing leaves the building without passing the same gates every time.

flowchart TD
  subgraph IN [Workflow: intake]
    direction LR
    E[Email arrives] --> C[AI: classify] --> R{Routing rules}
  end
  subgraph AG [Bounded agent: draft the reply]
    direction LR
    A{Next lookup?} --> O[Order system]
    A --> H[Help centre]
    A --> P[Past tickets]
  end
  subgraph OUT [Workflow: gates]
    direction LR
    Q{Policy checks} --> H2[Person approves] --> S[Send reply]
  end
  IN --> AG --> OUT

In code, the bounded step is small. This is a simplified sketch rather than a framework: the model call is passed in, and everything that makes it safe to run lives around it.

python
from dataclasses import dataclass

ALLOWED_TOOLS = {"lookup_order", "search_help_centre", "recent_tickets"}
MAX_STEPS = 6


@dataclass
class Draft:
    reply: str
    confidence: float  # 0.0 to 1.0, reported by the model and checked below
    sources: list[str]


def draft_reply(ticket, call_model, tools) -> Draft | None:
    """Let the model explore, but only with these tools and within these limits."""
    history = [ticket.as_prompt()]
    for _ in range(MAX_STEPS):
        action = call_model(history, tools=sorted(ALLOWED_TOOLS))
        if action.kind == "final":
            return Draft(**action.payload)  # fails loudly if the shape is wrong
        if action.tool not in ALLOWED_TOOLS:
            raise PermissionError(f"tool not allowed: {action.tool}")
        history.append(tools[action.tool](**action.args))
    return None  # out of steps: the workflow hands the ticket to a person

The thresholds that decide what happens next are business decisions, so they sit in configuration where the owning team can change them without a deploy:

yaml
support_reply:
  auto_approve_above: 0.92   # still logged and sampled for review
  human_review_below: 0.92
  never_auto_send:
    - refunds
    - legal
    - account_closure
  max_cost_per_ticket_usd: 0.05

A few rules make this pattern work:

  1. Keep the tool list short and explicit. An agent can only call what you allow, with the permissions of that step, not the whole system.
  2. Set hard limits. Cap the number of steps, the time and the spend per run, and define what happens when a limit is hit.
  3. Return structured output. Have the step return data in a fixed schema that the workflow validates, rather than free text that downstream code has to interpret.
  4. Use confidence to route. When a model is unsure, the workflow should send the case to a person instead of guessing. Thresholds are business decisions, so keep them in configuration.
  5. Put people at the irreversible points. Payments, customer-facing messages and record changes are natural places for human approval.

What production needs either way

Whether you build a workflow, an agent or a hybrid, the same engineering work separates a demo from a production system:

  • Evaluation. A set of real, representative cases with known good outcomes, run on every prompt or model change. Without it you cannot tell whether a change made things better or worse.
  • Tracing and observability. Every model call, tool call, input and output should be logged and linked to the run that produced it. For agents this is not optional; it is the only way to debug a path you did not write.
  • Guardrails at the boundaries. Validate inputs and outputs, restrict tool permissions and redact sensitive data before it reaches a model that should not see it.
  • Cost and latency budgets. Know what a run should cost and how long it should take, and alert when it drifts.
  • Fallbacks. Decide what happens when a model is slow, unavailable or wrong: retry, switch model, or hand off to a person.
  • Versioning. Treat prompts, model versions and configuration like code, so any result can be traced to exactly what produced it.
  • Loose coupling to models. Models change quickly. Keep them behind an interface so you can move between providers, or to an open model running in your own infrastructure, without rebuilding the integrations around them. We cover how we approach that integration layer on the technology page.

A quick decision checklist

Before building, answer these for the process in front of you:

  1. Can you describe the steps today, and are they mostly the same each time? If yes, start with a workflow.
  2. Does any single step require open-ended exploration across several tools? If yes, consider a bounded agent inside that step.
  3. Where are the irreversible actions? Put approvals there regardless of design.
  4. What does a wrong answer cost, and how will you detect one? That decides how much evaluation and review you need.
  5. What volume do you expect? High volume favours predictable paths and small, specialised models where they are good enough.

If most answers point to a workflow, build the workflow first and add agent behaviour only where a specific step proves it needs it. It is much easier to loosen a well-understood process than to tighten an agent that was never bounded.

Where to go from here

The right design depends on the process, the systems it touches and the constraints around it, which is why we start every engagement by mapping the workflow before choosing tools. You can read how that works on how we work, or get in touch to talk through a workflow of your own.

Bring the problem and the decision that is stuck.

In one hour we look at the workflow, the systems around it, and whether AI is the right tool, and leave you with a clear next step.

Issued for
Discovery
Duration
1 hour
Bring
The process, the systems, the stuck decision
Leave with
A clear next step