AI on the Clock mehmeterkek.com

Nº 005 · ENGINEERING · 11 August 2026

What is Harness in AI

Everyone is debating: “Which model is better?” But the same model can go from 38% to 66% success when you change the system around it. The difference isn’t the model — it’s the scaffolding built around it. It’s called the harness. 👇 — Agent = Model + Harness The model is the brain: it reads context and makes decisions. The harness is everything around it — tools, sandbox, memory, verification, guardrails, logging, and execution. A raw model can’t write a file, run your tests, remember what happened last week, or recover from a crash. All of that comes from what’s built around it. As LangChain’s Vivek Trivedy put it: “If you’re not the model, you’re the harness.”

AI on the Clock No. 005, Engineering — "The part nobody benchmarks: Harness." Definition: everything around the model — tools, loop, context, memory, sandbox, verification, guardrails. The model is the part you choose; the harness is the part you build.
The definition: Agent = Model + Harness. The model is the brain — it reasons and decides. The harness is the body and the workspace: it acts, remembers, verifies, and enforces limits. Quote from Vivek Trivedy of LangChain: "If you're not the model, you're the harness."
How it runs: the model decides an action, then an arrow points down to the harness, which executes the tool, captures the result, and updates context — repeating until done or stopped. This is the ReAct loop (Yao et al., 2022); the complexity isn't in the loop, it's in everything the loop manages.
The progression, each step leading to the next: 1) prompt engineering — wording the ask; 2) context engineering — curating what the model sees; 3) harness engineering — designing the whole system around it. Each contains the last: prompts and context live inside the harness.
The building blocks of a harness: system prompt (standing instructions and constraints); tools and execution (model decides, harness executes); sandbox and files (workspace where the agent operates); memory and compaction (what persists, what gets summarized); verification (test, inspect, validate, correct course); guardrails and observability (approvals, limits, logs, full trace).
The evidence: researchers fixed the model, then ran the same coding tasks through five different harnesses. GLM 5.1 went from 60.9% to 73.4%, a gain of 12.5 points; Qwen 3.6-flash went from 38.6% to 66.0%, a gain of 27.4 points. Only the harness changed, and a separate study of 106 tasks and 5,194 runs reached the same conclusion. Sources: CLAW-SWE-Bench and Harness-Bench (arXiv, 2026); pass@1 means solved on first attempt.
Where harnesses break: context rot — history grows and reasoning degrades; tool overload — too many options, worse decisions; brittle wiring — a reworded description breaks the call; weak verification — success declared on incomplete work; missing guardrails — irreversible action with no approval. These aren't necessarily model problems, but the model gets the blame anyway.
Time's up: model choice is a decision, harness design is a discipline. Three takeaways — start dumb with one reliable loop first; write tool descriptions like docs; instrument everything, because you can't debug what you can't see. Save this for your next agent build.

Download the deck (PDF, 8 pages) Read the original on LinkedIn

Everyone is debating: “Which model is better?” But the same model can go from 38% to 66% success when you change the system around it. The difference isn’t the model — it’s the scaffolding built around it. It’s called the harness.

Agent = Model + Harness.

The model is the brain: it reads context and decides. The harness is everything around it — tools, sandbox, memory, verification, guardrails, logging.

A raw model can't write a file, run your tests, remember last week, or recover from a crash. All of that comes from what's built around it.

As LangChain's Vivek Trivedy put it: "If you're not the model, you're the harness."

HOW IT RUNS

Model decides an action → harness executes the tool, captures the result, updates context → repeat until done or stopped.

That's the ReAct loop (Yao et al., 2022). The complexity isn't in the loop — it's in everything the loop manages.

THE PROGRESSION

→ Prompt engineering — wording the ask → Context engineering — curating what the model sees → Harness engineering — designing the whole system around it

Each contains the last.

THE EVIDENCE

Researchers fixed the model, then ran the same coding tasks through five different harnesses:

→ GLM 5.1: 60.9% → 73.4% solved on first attempt → Qwen 3.6-flash: 38.6% → 66.0%

Same model. Same tasks. Only the harness changed.

A second study (106 tasks, 5,194 runs) concluded: capability should be reported per model-harness pair, not per model.

WHERE HARNESSES BREAK

→ CONTEXT ROT — history grows, reasoning degrades → TOOL OVERLOAD — too many options, worse decisions → BRITTLE WIRING — a reworded tool description silently breaks the call → WEAK VERIFICATION — success declared on incomplete work → MISSING GUARDRAILS — irreversible actions, no approval

These aren't necessarily model problems. The model gets the blame anyway.

THREE RULES

  1. Start dumb — one reliable loop before planning layers or sub-agents.
  2. Write tool descriptions like documentation. The model reads them.
  3. Instrument everything. You can't debug what you can't see.

Model choice is a decision. Harness design is a discipline.

Before you swap models again, look at the scaffolding around the one you already have.

What's the worst agent failure you've debugged — and was it really the model? 👇

Sources: Claw-SWE-Bench · Harness-Bench (arXiv, 2026)

AIontheClock #AIAgents #AIEngineering #AgenticAI #HarnessEngineering