HarnessTax: How Much Does the Harness Matter for Coding Agents?
What does a coding-agent harness actually add, and at what cost? UC Berkeley Sky Lab evaluated 21 model–harness pairs — seven models across Claude Code, Codex CLI, and Pi on SWE-bench Lite and Terminal-Bench 2.0. The harness barely moves success rate (±2–5%) but can swing cost up to 5×. A minimal four-tool open-source harness matches frontier success at a fraction of the cost, and models frequently win outside their own provider's harness — your Claude models may not need Claude Code.








