A model on its own only predicts text. The harness runs the loop that sends it context, executes the tool calls it asks for, feeds the results back, and decides when to stop. It also owns the parts that make an agent dependable over long runs: what goes into the context window, where progress is written down, what the agent is allowed to touch, and which deterministic checks it has to pass.
Because the harness and the model always run together, evaluating an agent means evaluating both. Changing the harness, for example by redesigning a tool or adding an independent verifier, can move results as much as changing the model, which is why “harness engineering” has become a discipline of its own.
