hraness
Theme
Appearance

What is an agent harness?

by hraness · drafted with ai assistance

An agent harness is the software around a language model that turns a text predictor into something that can do work. It holds the instructions, offers the tools, runs the loop, keeps the session alive between turns and across restarts, and speaks to whichever model API is on the other end. Earendil’s “What is a Harness?” gives the compact version: a piece of software that provides an environment for an AI model to operate within. Hraness keeps a reading digest of that post. This essay separates what the harness owns from what the model decides, then covers how that split changed the way agents are measured, what fails when the line is drawn badly, and how ALGAL rebuilds the harness around programs you can store, replay, and evolve.

What the model lacks on its own

A model call is stateless. Tokens go in, tokens come out, and nothing remains. The model has no memory between calls, no hands, no clock, no files, and no way to know whether the last thing it said worked. Every property people attribute to an agent, such as remembering the task, running a command, reading the result, and trying again, is supplied from outside the weights. The Prime Agent paper from Prime Intellect states the constraint plainly: language models are sequential processors, and long-horizon agency requires external information and computation beyond model weights and active context.

That outside is the harness. Presenting Prime Agent at Y Combinator’s Harness Night, Seth Karten compared a raw model to a Turing machine reading a tape and a harnessed model to a von Neumann computer with memory it can read and write. The comparison puts the upgrade outside the model. Becoming an agent left the weights unchanged; what the model gained was a program around it that can hold state and act on the world.

What a harness owns

Earendil lists four jobs: instructions, tools, the agentic loop, and a translation layer. Running harnesses in production adds a fifth, the session. Together they mark the boundary between what the harness owns and what the model decides.

Context. The harness decides what the model sees. That starts with the system prompt and extends to instruction files such as AGENTS.md or CLAUDE.md, repository conventions, prior decisions, and whatever memory survives from earlier sessions. It also includes the policy for what leaves the window. An empirical study of harness design for coding agents (Fan et al., 2026) held one execution loop fixed across 176 settings on SWE-Bench Verified and Terminal-Bench 2.1 and found that context management matters most when the window is tight, mostly by preventing overflow failures, and that rule-based elision before any model summarization is the cheapest strong policy. Ryan Lopopolo’s Harness Engineering anthology treats the model as a black box and names context and tools as the two external levers left to the people who run it. Context is the first one.

Tools. The harness describes and implements the capabilities the model can call: read a file, run a shell command, search the web, open a browser. Earendil’s observation is that the harness usually does not dictate when or how a tool is used. It makes the tool available and describes it clearly, and the model chooses. The description is not free. HarnessTax, from Berkeley’s Sky Lab, measured the first model call across seven models and found that Claude Code’s mean initial context was more than 10 times Pi’s, through longer instructions and larger tool schemas. Pi reached the cost and success frontier on both benchmarks with four tools: read, write, edit, and bash.

The loop. The model assesses the request, acts through a tool, reads the result, and reassesses. The harness runs that cycle and owns its edges: how many turns are allowed, how much can be spent, what counts as done, what happens when a tool fails, and whether the loop wakes on a user message, a schedule, or its own initiative. Laude Institute’s Headlong shows how far the last choice can go. Its agent never sleeps; incoming messages land as observations in one continuous stream of thought, and the agent decides whether to reply. The model still chooses each next action. The harness decides that there is a next action at all.

The session. Work has to outlive a single call and, in practice, a single process. The harness keeps the transcript, compacts it when it grows too large, checkpoints state, recovers after a crash, and coordinates subagents that share the work. Prime Agent backs each session with a persistent daemon and describes the harness as a membrane that standardizes execution, recovery, verification, and accounting so that a model fails because the task exceeded it, not because the harness dropped state. That sentence doubles as a definition of harness quality.

Translation. Every provider exposes a different API for messages, tool calls, and streaming. The harness normalizes them so the same session can run against a model from Anthropic, OpenAI, or an open-weight provider. Earendil frames this as the layer that moves control to the user: the harness is where your sessions, your tools, and your cost comparisons live, and the model is the part you can swap.

Against those five, the model owns one thing: judgment inside a turn. Which tool to call, what to write, whether the evidence is enough. A good harness keeps that boundary sharp. A weak one lets the two blur, and then nobody can say which side failed.

Why the noun matters

Claude Code shipped as an application for coding with one provider’s models. Codex followed the same shape. Cursor put a harness inside an editor. LangGraph made the loop a framework, and DeepSeek Harness made every capability a plugin. Pi and Headlong went the other way and made the whole thing small enough to read end to end. The reason all of these are now called by one name, and the noun has settled quickly enough to earn a Wikipedia entry, is that the category forces an ownership question. If the model is the product, your session lives in the lab’s app and your tools are whatever the lab shipped. If the harness is the product, the session, the tools, the memory, and the translation layer are yours, and the model becomes a replaceable engine. Earendil’s line is that unlike the models themselves, you can own and adapt the harness.

The second reason is measurement. Once a harness is a variable, a benchmark row that names only the model is underspecified. Two 2026 results show why, and they look contradictory until the boundary above is applied. HarnessTax ran 21 model and harness pairs, seven models across Claude Code, Codex CLI, and Pi, on SWE-bench Lite and Terminal-Bench 2.0. Harness choice moved success rates within a few points, but the same model at a similar success rate cost up to five times more in one harness than another, and in nine of twelve comparisons across Anthropic and OpenAI models, a harness other than the provider’s own reached the highest observed success rate. Prime Agent reports raising ARC-AGI-3 RHAE Best@1 from 30 percent to 95.5 percent with a harness built around a persistent REPL, programmatic context, and recursive subagents. Y Combinator opened Harness Night with that gap, attributed to the harness alone: same weights, better harness.

The two findings fit. On tasks the model can already complete in a few dozen turns, the harness mostly changes cost, and a large one imposes a tax on every call. On tasks that need state the model cannot hold, computation it cannot do in its head, and horizons longer than one context window, the harness changes whether the task is possible at all. Either way the harness belongs on the row.

What fails without one

With no harness, you are the harness. You paste tool output into a chat, you remember what was decided, you notice when the window fills, and the transcript belongs to the vendor. That works for a question and stops working at the first task that outlasts your attention.

A weak harness fails in recognizable ways. The window overflows and the run dies mid-task, which is the failure Fan et al. found context management mostly prevents. The process exits and the work vanishes with it, because the session lived in memory. A tool call cannot be replayed, so when the outcome is wrong nobody can tell whether the model chose badly or the harness dropped a result. A tool runs with whatever permissions the host process has, so the agent can reach anything the person running it can, and privileged information leaks into contexts where it does not belong; YC’s QM team names social context and permissions as the gap their own harness still has. The loop gives up early, which QM counters with grind budgets that forbid stopping before a set spend or wall-clock time. The loop stops itself: Headlong’s agent shut down its own service three times by accident before Laude added a guard that refuses the attempt. The harness trains the model without meaning to: Headlong’s 30-second inactivity watchdog killed spawned copies while they were thinking, and after 40 minutes of fighting it the agent mostly stopped delegating. And an oversized harness taxes every call before the model has read the task.

A better model would fix none of these, because each one sits on the harness’s side of the line.

Where harnesses are going

Francois Chaubard’s history at Harness Night splits the field into two eras. The static era added capabilities one at a time: few-shot examples in the prompt, chain of thought, tool calls, a memory the model could edit, skills distilled from successful tool sequences, subagents, and recursive calls. The harness got more capable and stayed fixed. The self-improving era, the last six months by his count, lets the harness change: DSPy searches over system prompts, Darwin-Gödel machines edit the harness code itself, and Prime Agent’s Continual Harness keeps histories, memories, skills, prompts, and subagent specifications across runs. Once a harness can propose changes to itself, the open problem is deciding which changes to accept, measuring whether they helped, and keeping a record a person can check.

How ALGAL reimagines the harness

Every harness above is a program written in a general-purpose language, from TypeScript and Python to Rust and, in Headlong’s case, Bash. The loop, the tool calls, and the session state are wherever the author put them. State is mutable and ambient, branching is an if statement or a sentence in the prompt, the tool log is whatever got printed, and the evidence of a run is a transcript. That is fine for a human reading along. It is a poor foundation for a harness that is supposed to improve itself, because nothing about it can be held, compared, or verified as a whole.

ALGAL treats the harness itself as a value, a program that can be stored, compared, and verified as a whole. An ALGAL program is a typed, bounded, content-addressed graph. Its cells are inputs, pure functions, tool calls, and model effects; its edges carry typed values; and a budget covers the whole run. Because the program is data with a digest, it can be stored, diffed, bundled, sent to another host, or spawned as a child of another program, and two hosts can agree they are running the same thing. A native Rust kernel and a Bun runtime execute the same manifest, and a program paused in one can resume in the other.

The model’s role is declared in the structure rather than described in prose. decide and generate are effects with an explicit context view, a typed output contract, and a place in a budget that counts executor attempts. A decision selects one arm of an exhaustive match and a classifier’s label drives guarded edges, so routing lives in the graph where it can be inspected and replayed. Where an ordinary harness asks the model to plan, act, and judge in one open loop, ALGAL keeps that open loop outside and hands it a library of bounded programs it can call with confidence. The outer loop decides what to do; the program decides how, within limits it cannot exceed.

Every run emits a receipt bound to the program’s digest. Pure work is recomputed on verification and effects replay from the record, so a reviewer can check a run offline without the original store, the credentials, or the model. This is the membrane Prime Agent describes, made checkable. When an outcome is wrong, the receipt separates what the model judged from what the runtime did. The boundary of that claim matters: a verifying receipt proves the execution was consistent with the program, not that a model’s answer was true.

Authority is bounded the same way. A manifest carries no host code and can only name capabilities the host has already admitted, meaning registered and allowed. Each tool declares a read or write class, byte and time limits, and an idempotency key, and the host, not the program, decides which tools exist. A human approval is a wait in the structure: the first command exits, the proposal is retained, and a later command authorizes that exact report. The model cannot approve itself. Those two properties together are what make an untrusted program safe to run, which is the prerequisite for letting programs write programs.

That is where ALGAL leaves the existing category. Because programs are values and receipts are evidence, a program can propose a child program as data, and the host can admit it and run it under the parent’s bounds. A shared store with a tool and function registry becomes a habitat: a population of programs that evolve through search and host admission. Evaluation runs candidates against declared cases and keeps their receipts; a host-controlled selection policy decides what is promoted; and lineage becomes a record in which every proposal, measurement, and promotion is addressed by its content. A minimal harness like Pi, in this view, is a seed program in a habitat, and a habitat can grow into a swarm of agents that share artifacts and tools. Proposal alone establishes nothing. A child that ran successfully has not shown that it improved the task until the host’s cases say so.

ALGAL is working prerelease software for bounded workflows that the host controls. It is an application VM, not an operating-system sandbox, so host tools keep their host permissions. Its packaged demos use labeled deterministic decision fixtures: the durable state, journals, recovery, and offline verification are real, and the demos do not evaluate live model quality. The habitat example is a deterministic end-to-end prototype and the civilization loop is a first runnable version; neither is a finished product.

Definition

An agent harness is the program around the model: the context it sees, the tools it can call, the loop that gives it a next turn, the session that lets work outlive a call, and the translation layer that makes the model replaceable. Earendil’s climbing-kit argument is that you should own that program. ALGAL’s extension is that when the program is a value with a receipt, you also own its history, you can prove what it did without asking the model again, and you can let it propose its successors while keeping the decision to admit them.