Graph, Harnesses and Loop Engineering Explained

In this article:
Subscribe to our blog:

When a coding agent fails, the usual instinct is to swap the model. The evidence points somewhere else. On SWE-bench Pro, Claude Opus 4.5 scored 45.9% in Scale AI’s standardized test harness and 55.4% running inside Claude Code. That’s a 9.5-point difference from the setup alone, when a model upgrade typically moves scores by 2 to 4 points.

That setup is what graph, harness, and loop engineering describe. The three terms arrived within months of each other, and there’s still no consensus on where one ends and the next begins. One 2026 study puts the decision to retry or stop inside the harness. Another puts the loop above it.

In this article, we’ll cover how we draw the lines between what each layer owns, how the three fit together, how to tell which one is failing, and the step many teams skip before any of it.

TL;DR: what’s the difference between graph, harness, and loop engineering?

Graph engineering decides in what order work happens and where it goes after each step. Harness engineering decides what the agent can touch, from tools to permissions, and what checks its output. Loop engineering decides when the agent retries, escalates, or stops.

Together, they’re three decisions within one system, so when an agent fails, engineering teams can fix the layer that owns the failing decision rather than just changing the model.

What is graph engineering?

Graph engineering decides in what order work happens and where it goes after each step. Instead of packing a workflow into one long prompt, you lay it out as nodes and edges:

  • Nodes do the work. A node can be an agent, a deterministic step like a build, a validator, or a human approval.
  • Edges decide what happens next, either always or based on what the previous node returned.

This matters because agent work rarely runs in a straight line. Say an agent is building a data pipeline that scrapes public websites, cleans the data, and stores it.

The scraper, cleaner, and storage code can be built in parallel. Each piece needs its own checks, failures have to go back to the step that caused them, and nothing should touch live sites until a person signs off. A single prompt can’t express any of that.

Three patterns cover most cases:

  • Sequential stages for steps with a fixed order, like running the build before the tests.
  • Conditional routing for branching on a result, like sending a failed security check back to the agent while a passing one moves on to human approval.
  • Orchestrator-worker for work that splits in ways you can’t define upfront, like a planning agent dividing the pipeline into modules and handing each to a separate agent.

When an agent runs steps out of order or loses track of earlier decisions, the graph is usually the layer to fix.

Long context makes it worse: Chroma’s 2025 study of 18 models found performance dropped as input length grew, even on simple tasks. Making state explicit at each node keeps every step working from a short, relevant context.

What is harness engineering?

Harness engineering decides what the agent can touch, from tools to permissions, and what checks its output. In practice, the harness covers:

  • Tools: the commands, APIs, and services the agent can call.
  • Sandbox: where the agent’s code runs, isolated from everything else.
  • Permissions: which files, repositories, and systems it can read or change.
  • Saved state: what survives a crash, so the agent can pick up where it left off.
  • Checks: the tests, linters, and analysis its output has to pass.

Part of every harness comes built into the agent. Birgitta Böckeler calls the rest the outer harness: the rules, checks, and constraints your team adds for its own codebase. That’s the part you own.

Coming back to our data pipeline example: the harness is why the scraper can only fetch from approved domains and the storage step can write to a staging database but not to production. It’s also why all the code runs in a sandbox.

Running our own agents, we learned that if a check is optional, the agent skips it. A tool the agent can choose to call is a request. A hook that runs automatically after every edit is a guarantee.

Rules that live only in people’s heads get expensive. OpenAI’s Codex team spent every Friday, 20% of their week, cleaning up “AI slop” until they encoded their standards in the repository and automated the cleanup.

When an agent reaches systems it shouldn’t, ships output nothing has checked, or can’t resume after a crash, the harness is the layer to fix.

Note: We go deeper on what a strong harness should check in our guide to AI agent harnesses.

What is loop engineering?

Loop engineering decides when the agent retries, escalates, or stops. Those decisions live in the code around the model, not in the model itself, which is why a better model won’t fix a badly designed loop.

A loop is the simplest kind of graph: one node that sends work back to itself until it’s done. Each pass follows the same cycle: the agent acts, something checks the result, and the loop decides what happens next.

A 2026 review of industry writing on loop engineering found broad agreement on what a production loop needs, including:

  • A cap on attempts, so a stuck agent doesn’t burn tokens and time without anyone noticing.
  • A definition of “done” that something outside the agent can check. The harness provides the check, and the loop decides whether to retry, roll back, or stop based on the result.
  • A point where a human takes over once the agent stops making progress.

Setting the cap is a judgment call. When we ran our own agents, retries beyond about two rounds stopped improving the code and started creating new problems as the agent overcorrected its own fixes. Capping self-correction at two passes and handing the rest to a person worked better.

When an agent keeps retrying without getting closer to done, the loop is the layer to fix. Tighten the stopping conditions or redefine “done” before you reach for a different model.

Note: For a loop running in production, see how Black Box runs PR review gates at scale.

How graph, harness, and loop engineering fit together

In a real run, the three layers work at the same time. Take an agent building a data pipeline. One step writes a scraper that collects data from public websites, another writes the code that cleans that data, and a third writes the code that stores it. Each layer handles a different part of that run:

  • The harness sets the boundaries every step works inside, such as which websites the scraper can reach and which database the storage step can write to.
  • The graph moves work between the steps, through a quality and security check, and on to a person for approval.
  • Loops run inside each step until its code passes its tests.

The layers also hand work to each other. When the quality and security check fails, the graph sends the work back to the step that caused it, and that step’s loop starts again. LangChain sees the same pattern in production: retries, requests for missing information, and pauses for human input all send work back to an earlier step.

HARNESS · what the agent can touch HTTP fetcher · allowlisted domains only Python sandbox · no prod secrets Staging DB: write · Prod DB: none Test runner + sample datasets External code checks (e.g. Codacy) Saved state · resume after a crash pass approved fail → back to the step that owns it AGENTPlan & splitinto 3 modules AGENTBuild scraperfetch + parse code AGENTBuild cleanerdedupe, normalize AGENTBuild storageschema + load jobs CHECKQuality & securitylint, scan, unit tests SCRIPTEnd-to-end testfull run on sample HUMANApprovebefore live sites SCRIPTRun pipelinescrape→clean→store retry · max 2 retry · max 2 retry · max 2 HARNESS · what the agent can touch HTTP fetcher · allowlisted domains only Python sandbox · no prod secrets Staging DB: write · Prod DB: none Test runner + sample datasets External code checks (e.g. Codacy) Saved state · resume after a crash IN PARALLELfailpassapproved AGENTPlan & splitinto 3 modules AGENTBuild scraperfetch + parse code AGENTBuild cleanerdedupe, normalize AGENTBuild storageschema + load jobs CHECKQuality & securitylint, scan, unit tests SCRIPTEnd-to-end testfull run on sample HUMANApprovebefore live sites SCRIPTRun pipelinescrape→clean→store retry · max 2 retry · max 2 retry · max 2

One piece connects all three layers: the check. The harness runs it, the loop decides whether to retry based on the result, and the graph routes work depending on whether it passed. If the check is weak or missing, all three layers make decisions on bad information.

Because each layer affects a different share of the system, we recommend treating changes to each one differently:

  • Treat harness changes like infrastructure changes. Giving the storage step write access to production or adding a new approved website affects every agent at once, so review it carefully. OpenAI’s Codex team describes the same approach: “enforce boundaries centrally, allow autonomy locally”.
  • Treat graph changes like process changes. Adding a human approval before the scraper hits a new website affects one workflow, so revisit the graph when the process itself changes.
  • Treat loop changes like tuning. Raising the cleaning step’s retry limit or tightening what “done” means affects a single step, so let the people closest to the task adjust it.

That separation is what makes a failed run traceable. When something breaks, you can ask which layer made the wrong decision and fix that layer, instead of guessing or swapping the model.

Which layer is failing? A diagnostic playbook

When a coding agent fails, the instinct is to call the whole agent broken. In our model, graph, harness, and loop engineering each own a different decision, so each layer fails in a recognizable way.

Symptom

Layer

Typical fix

The agent runs steps out of order or loses track of earlier decisions mid-task

Graph

Make the order of steps and the state passed between them explicit

The agent reaches systems it shouldn’t, ships output nothing has checked, or can’t resume after a crash

Harness

Restrict permissions, add an external check, and save state between steps

The agent keeps retrying without getting closer to done

Loop

Cap attempts, tighten stopping conditions, or redefine “done”

If you can’t tell which layer is at fault, look at the fix you need:

  • A change to the order of steps or how work is handed off belongs to the graph.
  • A permission change belongs to the harness.
  • A new stopping condition or definition of “done” belongs to the loop.

Once you’ve found the layer, make the fix permanent. Every repeated mistake should become a rule: a lint check, a line in the agent’s instruction file, or a new gate. That way, the same failure can’t come back unnoticed.

Ask which layer owns the failure, not which model to switch. Then name an owner for that layer, so the next time the same failure shows up, someone is already responsible for it.

Why agent engineering starts with an AI inventory

Graph, harness, and loop engineering all assume you know which coding agents are running in your repositories. Many engineering leaders can’t answer that question with confidence, and you can’t set boundaries, routes, or stopping conditions for an agent you haven’t found.

The problem usually isn’t chaos. When we scanned more than 6,700 repositories across more than 800 organizations, we found 27 different coding assistants. Claude Code appeared most often, followed by Cursor and Copilot. But the average organization showed traces of only two, so adoption is limited but rarely measured.

Knowing which agents exist is only the first step. A useful inventory also records what each agent is allowed to do: its tools, credentials, and approval gates. OWASP’s Agent Observability Standard calls this an Agent Bill of Materials (AgBOM), and it’s effectively a record of each agent’s harness.

That record has to show what agents actually did, too. Researchers studying loop engineering mined 36,710 open-source repositories. Teams committed the configuration for their agent loops but almost never the state files that record what each run did. A stopping condition is only as good as your ability to confirm it fired.

An inventory also can’t be a one-time audit. An agent’s permissions can expand through a single config edit, without any code changing, and developers adopt new tools all the time. Tracking changes continuously keeps the inventory accurate enough to act on.

Where Codacy fits across graph, harness, and loop engineering

Codacy doesn’t cover the three layers equally, and it isn’t meant to. Our role is the check every layer depends on, plus the inventory that tells you which agents need checking in the first place.

  • Inventory: Codacy’s AI Inventory keeps a continuously updated record of the coding tools, AI models, and API keys across every repository, so you know which agents you’re engineering for.
  • Graph: Codacy’s checks plug into the graph your team already runs as the validation step that decides whether work moves forward or goes back.
  • Harness: This is where Codacy works most directly. The Codacy MCP Server and Codacy Analysis CLI let an agent run Codacy’s analysis locally. The agent gets each issue’s exact file and line, along with duplication and complexity results. The Codacy Cloud CLI tells the agent whether a pull request passes its coverage gate.
  • Loop: Verity runs when an agent stops, and a different model, backed by Codacy’s deterministic analysis, reviews the change. If the change fails, the agent gets two attempts to fix it before the work goes to a person.

Every layer depends on that same check, so it’s the one part of the system that shouldn’t rely on the agent’s own opinion. Agent output gets checked against real rules, not graded on confidence.

Own the layer, not the model

Graph, harness, and loop engineering aren’t rival approaches competing for the same budget line. They’re three decisions inside one system: what order work happens in, what the agent can touch and what checks its output, and when it retries or stops.

Separating them turns “the agent is broken” into a fix someone can own. Swapping the model becomes the last resort, not the first.

Engineering the agent is now part of engineering code. Start by knowing which agents are running, then decide which layer owns each failure.

Start with the agents you already run

Codacy’s AI Inventory keeps a live record of the coding tools, AI models, and API keys across every repository. See which agents you’re working with before you decide what each layer should own.

Scan your repository for free →

Subscribe to our blog

Stay updated with our monthly newsletter.