Field notes

How to build reliable AI agent workflows with open-source Kortix

Map AI agent workflow failure modes to concrete controls, with a reference implementation on open-source Kortix.

The Kortix Team
Kortix··6 min read

Reliable AI agent workflows fail in ways a generic retry does not fix: a tool reports success while returning an error payload, an upstream schema drifts, or a loop has no exit. Treat reliability as engineering around the model. Name each failure mode, attach one control, and instrument the signal that proves the control holds. Kortix is the open-source platform that supplies those primitives, from a sandbox per session to a review gate on every change.

Why agent workflows fail in production

Agent workflow failures cluster into a small set of recurring modes. Each mode needs a different control and a different signal, so a generic "add retries" fix covers one row and leaves the rest.

Failure modeWhat it looks likeControlSignal
Silent tool failureA tool returns success with an error payload, an empty result or a stale valueValidate the result content, not the transport statusTool-error and empty-result rate
Schema driftAn upstream API renames or retypes a field while the agent parses the old shapeValidate typed input and output at every step boundarySchema-validation failures per step
Non-determinismThe same input produces different plans or outputs across runsPin the model and version, and run a fixed eval setEval pass-rate variance
Context rotThe model misses a fact as the context fills with distractorsRetrieve only relevant context, compress it, restate the taskTask-success rate against context size
Compounding errorsOne wrong step becomes the input to the nextProgrammatic gates between steps and a bounded verifier passGate-failure and rework-loop rate
Runaway loopsThe agent iterates because no exit condition ever evaluates trueCap iterations and set explicit stopping conditions and budgetsTimeouts and cost per task
Unguarded side effectsThe agent sends, deletes or charges without reviewHuman approval before irreversible actions, and least-privilege credentialsUnapproved actions, which should be zero
Duplicate side effectsA retried call creates a second resourceIdempotency keys, timeouts, capped backoff with jitter, a circuit breakerDuplicate-action and circuit-open counts

Context rot shows why the control has to match the mode: Chroma's evaluation found model performance degrading as input length grows, with topically related distractors reducing it further (Chroma). The structural choice comes before any of these controls, though: how much of the flow the model is allowed to decide.

Choose the simplest structure that fits

Anthropic separates a workflow from an agent (Anthropic). A workflow orchestrates models and tools through predefined code paths. An agent lets the model direct its own process and tool use. Both are agentic systems, and they carry different failure profiles: a workflow is predictable and rigid, while an agent is flexible and exposed to compounding errors.

Google Cloud reaches the same conclusion from the design side: start with a single agent so you can refine the core logic, prompt and tool definitions before adding more (Google Cloud). A single agent's performance drops as it takes on more tools and more complex tasks, which shows up as latency, wrong tool selection or unfinished work. Anthropic's advice is to find the simplest solution possible and add complexity only when it earns its place.

Bound the loop and make retries safe

The loop is where an agent's flexibility turns expensive. Google Cloud describes a loop pattern that repeats subagents until a termination condition is met, and warns that its "primary trade-off is the risk of an infinite loop" when that condition is not correctly defined (Google Cloud). Set an exit condition the workflow can actually evaluate, plus a maximum number of iterations and a budget for time and tokens. Anthropic lists stopping conditions such as a maximum number of iterations as part of the design (Anthropic).

Retries help with transient faults only when the operation is idempotent. AWS's guidance is that the effect of a retried call should happen once even when the call is made multiple times (AWS). For timing, capped exponential backoff with jitter spreads retries so clients do not retry in synchronized bursts (AWS). A timeout bounds each attempt, and a circuit breaker stops calling a service after failures cross a threshold so it can recover (Azure). For a long run, a durable execution system runs the workflow to completion and, after a crash, resumes at the point where it stopped (Temporal).

Validate every boundary and gate every side effect

Every boundary between steps is a place where a wrong shape enters the workflow. Define the expected input and output of each step and validate it against a schema before the next step consumes the result. JSON Schema is a common way to express that shape, and a mismatch fails fast instead of flowing downstream (JSON Schema).

Anthropic warns that agents carry "the potential for compounding errors", so a wrong step should not quietly become the input to the next (Anthropic). Put programmatic checks between steps and a bounded verifier pass with explicit pass criteria. Google Cloud's review-and-critique pattern has a critic evaluate output against predefined criteria, and Anthropic's evaluator-optimizer runs a generator and a critic in a loop (Google Cloud, Anthropic).

Some actions should never run unattended. Google Cloud asks whether a task involves high-stakes decisions, safety-critical operations or subjective approvals that need human judgment (Google Cloud), and Anthropic notes that agents can pause for human feedback at checkpoints or blockers (Anthropic). Gate irreversible actions behind approval, and give the agent only the credentials it needs.

Track the signals that prove a control holds

Observability has to cover both the deterministic and the probabilistic parts of a workflow. The OpenTelemetry GenAI semantic conventions define spans, metrics and events for GenAI clients and MCP, which gives traces a shared vocabulary instead of a one-off log format (OpenTelemetry).

Read those signals per step. Emit one span for each step with its inputs, outputs, latency and errors. Track the tool-error and empty-result rate for each tool, schema-validation failures, task-success and rework-loop rates, and cost per task. Hold a fixed eval set of representative tasks and run it before any change to a prompt, a tool or a model; Anthropic recommends comprehensive evaluation and sandboxed testing with guardrails (Anthropic).

Reliable workflows on open-source Kortix

Kortix is the open-source AI Management System: agents, shared skills, company memory and connectors live in one git repo you own. A session runs an agent in an isolated sandbox on its own branch, and the agent opens a change request a human reviews and merges (Kortix docs). Per-agent grants scope connectors, secrets and skills, connector credentials are brokered server-side so raw keys never enter the machine, and a full audit trail records what ran (Kortix on GitHub). Kortix is open source (Elastic License 2.0): self-host, read and modify the code.

# kortix.yaml: governance only; agent behavior lives in .kortix/opencode/agents/
kortix_version: 2
default_agent: workflow-runner
agents:
  workflow-runner:
    connectors: [github]
    secrets: []
    skills: [reliability-checks]
triggers:
  - slug: nightly-run
    type: cron
    agent: workflow-runner
    prompt: |
      Validate every tool result against its schema, stop after five
      iterations, and open a change request instead of writing to main.

The platform supplies the controls that are expensive to build twice: isolation per session, a review gate before work lands, scoped credentials and an audit trail. Self-hosting runs on a laptop, a VPS, your VPC or on-prem (Kortix on GitHub). You still define the schemas, exit conditions, budgets and eval set for each workflow.

If tool access is the part you are hardening, how to give AI agents tool access safely covers scoped connectors and approval policies. Open-source AI agent platforms compares the options on licence, deployment and model choice, and getting started with Kortix walks through the first session.

Start with one workflow. Name its failure modes, attach a control to each, and measure the signal that proves the control holds. Get started with open-source Kortix, or read the docs to run the reference shape yourself.

Run your whole company from one repo you own.

Start with one job and grow from there.

Get started