A harness is everything that decides what the model can see, do, remember and be checked against. It is where an agent stops being a prompt and becomes an engineering system with contracts, permissions and a replayable history.
Explore this article visually3 figures
01
A valid argument can still be forbidden
02
Same model, two behaviours
Our Mistral Vibe extension added three things to the CLI: reusable agent skills (a skill being a scoped instruction set plus the tools it is allowed to use), browser automation as a typed tool, and a router that sends a task to a local model or a cloud model depending on the task class. The model weights did not change. The completion rate on our test tasks changed a lot.
The unmodified CLI gave the model a broad shell and a long system prompt. With skills, the same model got a narrow tool surface per task and a shorter, specific instruction. It stopped exploring, because there was less to explore. That was the whole trick, and it is the trick behind every reliable agent I have seen: reduce the degrees of freedom until the remaining ones are the task.
Diagram 01
Five rings, one model
03
The five rings
Figure 02Data flow
A tool call crosses four contracts
Read every explanation
- Arguments
- Validate types and constraints before execution. A structurally valid path still needs to be checked against the permitted workspace.
- Authority
- Determine whether this caller may perform this action now. Neither the model nor retrieved content can grant itself additional authority.
- Execution
- Apply time and resource limits. For writes, decide how duplicate requests are detected before adding automatic retries.
- Result
- Return a bounded structured result and record provenance. The next step should not need to infer success from arbitrary console output.
Validate types and constraints before execution. A structurally valid path still needs to be checked against the permitted workspace.
I think of the harness as five concentric layers. The order matters because each layer is allowed to know about the one inside it and nothing else.
- Model. Weights, sampling settings, reasoning mode. The part everyone talks about and the part that changes least often in a working system.
- Context. Which files, documents or records the model is shown, chosen by a rule with a token budget. In a coding task this is the diff plus the files it touches, not the repository.
- Typed tools. Each tool has a JSON schema for its arguments and—more importantly—for its result.
run_testsreturns{passed, failed, first_failure}, not a wall of stdout. - Permissions. An allowlist of commands, a sandbox for anything that writes, and an approval gate for anything outside the working directory. The model does not know this ring exists; it just sees some calls refused with a reason.
- Trace and evaluation. Every call in and out becomes a span. A set of recorded runs becomes the regression suite that the next change is measured against.
{
"name": "run_tests",
"input_schema": { "type": "object", "properties": { "path": { "type": "string" } }, "required": ["path"] },
"result_schema": {
"type": "object",
"properties": {
"passed": { "type": "integer" },
"failed": { "type": "integer" },
"first_failure": { "type": ["string", "null"], "maxLength": 1200 }
},
"required": ["passed", "failed"]
}
}The result_schema is the part most tool definitions skip. Without it, the model receives whatever the tool printed and has to infer structure from prose, which is exactly the fragile step the harness exists to remove. Truncating first_failure to 1 200 characters is deliberate: a 40 KB stack trace is not information, it is context pollution.
04
Routing is a harness decision, not a model decision
The router in our Vibe extension chose between a local model and a cloud model by task class, never by "how hard does this look". Rewriting a docstring, renaming a symbol and summarising a diff went local. Anything that required planning across files or generating tests went to the cloud. The classification was a short lookup, not another model call.
| Task class | Route | Why |
|---|---|---|
| Single-file edit, explicit instruction | Local | Latency matters more than reasoning; failure is cheap and visible |
| Multi-file change, tests required | Cloud | Planning quality dominates; a wrong plan costs minutes |
| Browser automation step | Cloud + typed tool | The tool constrains the action; the model only fills arguments |
| Anything touching secrets or CI config | Refused → approval | Not a model question at all |
Putting routing in the harness means it can be changed, tested and audited without touching prompts. It also means the cheap path stays cheap: the local model never sees the long context the cloud path needs.
05
Traces are the API to your own past
Figure 03Comparison
Matched replay versus a diverging trajectory
Read every explanation
- Matched call
- The same tool, arguments and relevant state can use the recorded fixture for a controlled comparison of downstream behaviour.
- Changed call
- Do not return result A for a request to read B. Mark the divergence and evaluate the new trajectory in a separate controlled environment.
The same tool, arguments and relevant state can use the recorded fixture for a controlled comparison of downstream behaviour.
flightrec exists because I kept asking "what did it do last time" and having no answer. It records a run as a sequence of events—model call, tool call, tool result, state change, approval—and can replay that sequence against a new configuration. The replay does not re-execute side effects; it feeds the recorded tool results back and lets the new model choose its next step, so the two trajectories can be diffed.
The questions that become cheap to answer: did the new prompt change which files got read? Did the model upgrade add tool calls, or remove them? Did the permission change block something a run used to rely on? Each of those used to be a guess. A diff of two event logs is not.
06
Start with one task and five recorded runs
None of this needs a platform. One task, three typed tools, an allowlist, a trace file and five recorded runs is enough to find out whether a failure lives in context, tools, permissions or the model—and that diagnosis is what the harness is for. Models will keep changing. The rings around them are the part a team actually owns.
07
Replay stops being comparable when the inputs diverge
A recorded tool result is valid only for the tool name, arguments and state that produced it. If a new model requests another file or changes a query, replay must flag that mismatch rather than return the next result in the log. Use matched fixtures for controlled comparisons and a separate sandbox run to evaluate new trajectories. For retries of real writes, record an operation identifier and check whether the effect already happened before executing it again.
More articles
09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read