Harness Engineering for Reliable AI Agents

Harness engineering turns a language model into an agent that can finish useful work. The harness matters as much as the model because it controls tools, context, limits, and verification.
We use OpenCode for daily repository work. We also built agents/harness_sokrates, a small Python harness, to understand each moving part without a large framework.
What the harness controls
A model only returns text or requests a tool. The harness decides what happens next and sends the result back. A useful setup needs five parts:
- a clear system prompt and task,
- tools with narrow, readable contracts,
- an execution loop with firm limits,
- current repository and task context,
- checks that prove the result works.
OpenCode provides this wider working environment. It can inspect a codebase, edit files, run commands, track tasks, and delegate research. Its workflow also keeps human review visible.
What we learned from Sokrates
Sokrates strips the same idea down to a few modules. llm.py calls any OpenAI-compatible API. agent.py runs the tool loop. tools.py exposes HTTP fetches, shell commands, and file reads.
The loop stops after 12 steps. Tool output stops at 10,000 characters. These two limits prevent common runaway failures and protect the context window.
We tested the harness with a report agent. It reads a crawler report, visits the audited website, and writes advice in the site’s language. The task proves that a small agent can combine private data, public context, and strict output rules.
Small harness or complete platform?
| Need | Sokrates | OpenCode setup |
|---|---|---|
| Learn and test an agent loop | Best fit | More than needed |
| Work across a real repository | Basic tools | Full workflow |
| Track multi-step delivery | Planned | Built in |
| Delegate isolated research | Planned | Built in |
Sokrates still has roadmap items for file writing, live Git reminders, task tracking, context compaction, and sub-agents. We do not hide that gap. A production design should separate working controls from planned controls.
The practical result
Start with one business process and the smallest useful tool set. Add timeouts, output caps, logs, and an approval point before broader access.
Do not begin with a complex multi-agent graph. First prove that one agent can finish one measured task safely. Our Agentic AI engineering offer applies this process to business workflows. For model selection, see our review of models and agent harnesses.