A harness is everything that decides what the model can see, do, remember and be checked against. It is where an agent stops being a prompt and becomes an engineering system with contracts, permissions and a replayable history.

Explore this article visually3 figures

01

A valid argument can still be forbidden

02

Same model, two behaviours

Our Mistral Vibe extension added three things to the CLI: reusable agent skills (a skill being a scoped instruction set plus the tools it is allowed to use), browser automation as a typed tool, and a router that sends a task to a local model or a cloud model depending on the task class. The model weights did not change. The completion rate on our test tasks changed a lot.

The unmodified CLI gave the model a broad shell and a long system prompt. With skills, the same model got a narrow tool surface per task and a shorter, specific instruction. It stopped exploring, because there was less to explore. That was the whole trick, and it is the trick behind every reliable agent I have seen: reduce the degrees of freedom until the remaining ones are the task.

Diagram 01

Five rings, one model

model01 · modelcontext02 · contextselected files · token budgettyped tools03 · typed toolsvalidated schema · structured resultpermissions04 · permissionssandbox · allowlist · approval gatetrace + eval05 · trace + evalspans · replay · regression suitetrace → evaluation → tighter contractThe model sits at the centre; contracts and traces surround it.
Each ring only talks to its neighbours. The model sees selected context and typed tools; permissions and tracing wrap those without the model being aware of them. Swapping the core does not touch the outer rings.

03

The five rings

Figure 02Data flow

A tool call crosses four contracts

Select an element to explore its role.
Read every explanation
Arguments
Validate types and constraints before execution. A structurally valid path still needs to be checked against the permitted workspace.
Authority
Determine whether this caller may perform this action now. Neither the model nor retrieved content can grant itself additional authority.
Execution
Apply time and resource limits. For writes, decide how duplicate requests are detected before adding automatic retries.
Result
Return a bounded structured result and record provenance. The next step should not need to infer success from arbitrary console output.
01 / 04
Arguments

Validate types and constraints before execution. A structurally valid path still needs to be checked against the permitted workspace.

The application enforces these boundaries. Instructions in a prompt cannot replace validation or permissions.

I think of the harness as five concentric layers. The order matters because each layer is allowed to know about the one inside it and nothing else.

  1. Model. Weights, sampling settings, reasoning mode. The part everyone talks about and the part that changes least often in a working system.
  2. Context. Which files, documents or records the model is shown, chosen by a rule with a token budget. In a coding task this is the diff plus the files it touches, not the repository.
  3. Typed tools. Each tool has a JSON schema for its arguments and—more importantly—for its result. run_tests returns {passed, failed, first_failure}, not a wall of stdout.
  4. Permissions. An allowlist of commands, a sandbox for anything that writes, and an approval gate for anything outside the working directory. The model does not know this ring exists; it just sees some calls refused with a reason.
  5. Trace and evaluation. Every call in and out becomes a span. A set of recorded runs becomes the regression suite that the next change is measured against.
jsonA tool whose result is typed. The harness validates both directions; the model never parses stdout.
{
  "name": "run_tests",
  "input_schema": { "type": "object", "properties": { "path": { "type": "string" } }, "required": ["path"] },
  "result_schema": {
    "type": "object",
    "properties": {
      "passed": { "type": "integer" },
      "failed": { "type": "integer" },
      "first_failure": { "type": ["string", "null"], "maxLength": 1200 }
    },
    "required": ["passed", "failed"]
  }
}

The result_schema is the part most tool definitions skip. Without it, the model receives whatever the tool printed and has to infer structure from prose, which is exactly the fragile step the harness exists to remove. Truncating first_failure to 1 200 characters is deliberate: a 40 KB stack trace is not information, it is context pollution.

04

Routing is a harness decision, not a model decision

The router in our Vibe extension chose between a local model and a cloud model by task class, never by "how hard does this look". Rewriting a docstring, renaming a symbol and summarising a diff went local. Anything that required planning across files or generating tests went to the cloud. The classification was a short lookup, not another model call.

Task classRouteWhy
Single-file edit, explicit instructionLocalLatency matters more than reasoning; failure is cheap and visible
Multi-file change, tests requiredCloudPlanning quality dominates; a wrong plan costs minutes
Browser automation stepCloud + typed toolThe tool constrains the action; the model only fills arguments
Anything touching secrets or CI configRefused → approvalNot a model question at all

Putting routing in the harness means it can be changed, tested and audited without touching prompts. It also means the cheap path stays cheap: the local model never sees the long context the cloud path needs.

05

Traces are the API to your own past

Figure 03Comparison

Matched replay versus a diverging trajectory

Select an element to explore its role.
Read every explanation
Matched call
The same tool, arguments and relevant state can use the recorded fixture for a controlled comparison of downstream behaviour.
Changed call
Do not return result A for a request to read B. Mark the divergence and evaluate the new trajectory in a separate controlled environment.
01 / 02
Matched call

The same tool, arguments and relevant state can use the recorded fixture for a controlled comparison of downstream behaviour.

A conceptual comparison. Reusing the next log entry is valid only when its call and state match.

flightrec exists because I kept asking "what did it do last time" and having no answer. It records a run as a sequence of events—model call, tool call, tool result, state change, approval—and can replay that sequence against a new configuration. The replay does not re-execute side effects; it feeds the recorded tool results back and lets the new model choose its next step, so the two trajectories can be diffed.

The questions that become cheap to answer: did the new prompt change which files got read? Did the model upgrade add tool calls, or remove them? Did the permission change block something a run used to rely on? Each of those used to be a guess. A diff of two event logs is not.

06

Start with one task and five recorded runs

None of this needs a platform. One task, three typed tools, an allowlist, a trace file and five recorded runs is enough to find out whether a failure lives in context, tools, permissions or the model—and that diagnosis is what the harness is for. Models will keep changing. The rings around them are the part a team actually owns.

07

Replay stops being comparable when the inputs diverge

A recorded tool result is valid only for the tool name, arguments and state that produced it. If a new model requests another file or changes a query, replay must flag that mismatch rather than return the next result in the log. Use matched fixtures for controlled comparisons and a separate sandbox run to evaluate new trajectories. For retries of real writes, record an operation identifier and check whether the effect already happened before executing it again.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read