A single API call can produce an impressive answer. It can also approve the wrong refund, repeat the same tool call, leak sensitive context, or leave a system in an unknown state after a timeout.
That gap is where agent harness engineering begins.
The model is only one part of an agent. Around it sits a control system that decides what the model may see, which tools it may call, what those tools may change, how state is saved, when retries are safe, what gets recorded, and when the run must stop or ask for help.
This article explains that control system in plain English. It does not repeat a generic “model plus tools” definition or reproduce a Python architecture walkthrough. Instead, it focuses on the production questions that determine whether an agent is merely interesting in a demo or dependable enough to operate inside a real product.
The central idea is simple:
A model proposes. The harness disposes.
The model can suggest an action. The harness supplies the boundaries, executes only what is allowed, checks what happened, and owns the consequences.
Key takeaways
Click any topic to expand or collapseControlled Execution Loop
A bare LLM API call generates an output; a production agent needs a controlled loop around that output.
Retry & Timeout Hazards
Retries are not automatically safe. After a timeout, the side effect may have happened even when the client received no response.
External Security Boundary
Prompts are guidance, not authorization. Permissions, validation, isolation, and postcondition checks must live outside the model.
Observability vs Auditing
Good observability records the agent’s actions and outcomes without treating traces as hidden reasoning or a complete audit record.
Trajectory Evaluation
The best evaluation question is not only “Was the final answer good?” but also “Did the agent take the right path to get there?”
What is agent harness engineering?
There is no single industry-standard definition yet. The phrase is being used in overlapping ways: sometimes for the runtime that executes an agent, and sometimes for the broader development environment of tools, tests, repository rules, feedback loops, and operational controls around an agent.

A safe working definition is:
Agent harness engineering is the production discipline of designing the runtime and control system around a language model so the system can act within explicit capabilities, retain the right state, receive feedback, verify results, and recover or stop safely.
The harness usually covers:
- the agent loop and its terminal conditions;
- context assembly, memory, and state transitions;
- tool schemas, argument validation, and dispatch;
- identity, authorization, and approval;
- retries, timeouts, cancellation, and recovery;
- sandbox and network boundaries;
- structured traces, metrics, and audit events;
- evaluations that inspect both actions and outcomes.
OpenAI uses “harness engineering” in its discussion of the repository and feedback environment around Codex, while Microsoft documents an “Agent Harness” in its Agent Framework overview. Those usages are related but not identical. That is why the term should be explained as a working engineering boundary, not presented as a settled specification.
What the harness is not
The harness is not simply a system prompt, an SDK name, or a list of tools. It is not automatically safe because it runs inside a framework. It is not the same thing as memory, tracing, or a sandbox in isolation.
A useful distinction is this:
- The model generates text, structured output, or a proposed tool call.
- The prompt explains goals, rules, and formats.
- The tool layer validates and executes an action.
- The runner controls turns, budgets, handoffs, pauses, and termination.
- The runtime supplies identity, state, scheduling, storage, telemetry, and recovery.
- The harness is the engineering discipline that makes those controls work together.
Adding a model to a predefined workflow does not automatically make the workflow an agent. Anthropic’s practical distinction between workflows and agents is useful here: use a workflow when the path is known and controlled; use an agent when the system needs to choose its next actions dynamically.
| Layer | Responsible for | What it cannot prove |
|---|---|---|
| Model | Generating a response or proposed action | Authorization, execution, persistence, or verification |
| Prompt | Instructions, goals, constraints, and output format | IAM, isolation, or an enforceable security boundary |
| Tool layer | Validation, dispatch, permissions, result handling | That the downstream service honored every assumption |
| Runner | Turns, budgets, continuation, handoffs, and terminal states | Durable recovery or safe side effects by default |
| Harness | The combined control, execution, feedback, and evidence system | Universal reliability or a guarantee across every environment |
If the model can affect the world, the important design work is outside the model as much as inside it.
Why a bare API call is not a production agent
A minimal prototype often looks like this:
- Send a user request to a model.
- Receive text or a tool call.
- Execute the tool.
- Send the result back.
- Repeat until the model says it is finished.

That loop demonstrates the idea, but it leaves critical questions unanswered:
- What if the model never reaches a terminal state?
- What if the tool call is malformed or aimed at the wrong tenant?
- What if the request times out after the payment provider accepted it?
- What if a tool returns a malicious or misleading result?
- What if the process crashes between two steps?
- What if a subagent bypasses the parent agent’s guardrail?
- What evidence will an operator have after a customer reports a bad action?
The model does not own the answer to these questions. The harness does.
The difference is similar to the difference between a function and a service. A function can return the expected value in a test. A service must also handle timeouts, concurrency, permissions, retries, partial failure, logging, deployment, and recovery.
The same principle applies to agents. A successful demo proves that one path worked once. It does not prove that the system can handle an uncertain outcome safely.
The control loop: give the agent a reason to stop
An agent needs more than a goal. It needs a controlled execution loop.
The OpenAI Agents SDK runner places turns, tools, guardrails, handoffs, and sessions under one execution boundary, while LangChain’s middleware documentation covers model-call and tool-call limits. These are different implementations, but the shared engineering lesson is straightforward: the runtime must impose limits rather than wait for the model to decide when enough is enough.

At minimum, define:
- a maximum number of model and tool calls;
- a wall-clock deadline;
- a cost or token budget where the provider exposes the data;
- cancellation propagation;
- a no-progress or repeated-action detector;
- a policy path for escalation, pause, or human approval;
- explicit terminal states such as
completed,failed,blocked,needs_review, andunknown_outcome.
The final state matters. “The process stopped” is not enough. An operator needs to know whether the task completed, failed before the side effect, failed after it, or stopped without knowing what happened.
A practical run-state model
Do not store only the latest assistant message. Store the state that explains the run.
A useful run record contains:
- the user request and task identifier;
- the current step and allowed capabilities;
- proposed tool call and validated tool call;
- approval decision and approving identity;
- tool result or error class;
- side-effect status;
- retry count and reason;
- the next allowed transition;
- the final outcome and evidence.
This is where the harness becomes a control system rather than a chat loop.
A reliable agent is designed around terminal states, budgets, and transitions, not around the hope that the model will stop politely.
Guardrails: prompts guide behavior, policies enforce it
Guardrails are often discussed as if they were a single filter. In production, they are a set of controls placed at different trust boundaries.

A prompt can tell an agent, “Do not issue a refund above $500.” That is useful guidance. It is not an authorization boundary. A malicious instruction, a confused model, a tool bug, or a stale policy can still produce a dangerous request.
A stronger design checks the action in code or in a policy service before the side effect occurs:
- Is the user allowed to request this operation?
- Is the agent identity allowed to perform it?
- Is the target account or tenant correct?
- Are the arguments valid and within business limits?
- Does this action require approval?
- Has the exact action already been executed?
- Does the result satisfy the expected postcondition?
The OpenAI guardrails documentation distinguishes input, output, and tool-execution boundaries. The Google Agent Development Kit safety guidance covers identity, authorization, in-tool controls, callbacks, sandboxing, tracing, and evaluation. The OWASP AI Agent Security Cheat Sheet likewise treats tool access, least privilege, input validation, and output handling as security concerns rather than prompt-writing problems.
Put controls on both sides of the tool call
Before execution, validate the request. After execution, validate the result.
Pre-execution checks answer: “May this happen?”
Post-execution checks answer: “Did the expected thing actually happen?”
That second question is easy to miss. A payment tool can return a syntactically valid response with the wrong currency, account, or status. A CRM tool can report success while updating a different record than the one shown to the user. A postcondition check catches problems that a prompt cannot.
Guardrails must also follow delegation. A control attached to a parent agent is not automatically applied to every subagent, handoff, MCP server, or external API. Each trust boundary needs an explicit enforcement point.
Use prompts to steer. Use deterministic controls to enforce.
Tool contracts: make every action explicit before the model can call it
Many agent failures begin before the tool runs. The tool description is too vague, its arguments allow more than the business operation should allow, or the result does not say whether anything changed. A harness becomes easier to secure when every tool is treated as a typed contract rather than as a natural-language suggestion.
A production tool contract should answer six questions:
- What capability does this tool expose? Use a narrow verb and noun, such as
lookup_orderorrequest_refund, rather than a general-purposemanage_customer_accounttool. - Who may call it, for which tenant, and at what risk level? The answer must be enforceable outside the prompt.
- What arguments are required and what values are forbidden? Validate types, ranges, formats, ownership, and cross-field relationships before dispatch.
- Is the operation read-only, preparatory, or mutating? The label should determine whether approval, idempotency, or reconciliation is required.
- What does success mean? Define the postcondition the harness can verify, not merely the shape of the provider response.
- What happens when the result is delayed, duplicated, partial, or unknown? The contract needs an error and recovery path.
| Contract field | Weak version | Harness-grade version |
|---|---|---|
| Capability | Manage account | Read the order status for the authenticated customer |
| Arguments | Free-form customer data | Validated order ID bound to the current tenant and user |
| Effect label | Not specified | Read-only, no external mutation |
| Postcondition | Provider returned 200 | Returned order belongs to the requested customer and has a current status |
| Failure path | Try again | Classify the error, apply a bounded policy, and expose a safe result to the agent |
For mutating tools, add a dry-run or preview mode where the domain allows it. The agent can prepare the exact change, show the target and effect, and request approval before the commit step. This separates planning from mutation and makes human review meaningful: the reviewer approves a concrete operation, not a vague intention.
A tool description tells the model what it may propose. A tool contract tells the system what it may actually do.
Retries and idempotency: the side-effect problem most demos hide
Retries are easy when the operation is read-only. If a model request fails, retrying may cost more tokens but usually does not create a second customer order.
Tool calls are different.

Imagine an agent calls create_booking. The provider accepts the booking, but the network connection drops before the agent receives the response. The agent sees a timeout. Did the booking happen?
The correct answer is: unknown.
Retrying immediately can create a duplicate booking. Treating the timeout as a failure can lead an operator to repeat the action manually. Treating it as success can mislead the user. This is a distributed-systems problem before it is an AI problem.
AWS explains the pattern clearly in its guidance on making retries safe with idempotent APIs: use a caller-provided request identifier and reconcile the result when a timeout leaves the server outcome uncertain.
The three-part retry policy
A useful harness separates three decisions:
1. Is this error retryable?
A transient network error may be retryable. Invalid arguments are not. A permission denial is not fixed by repeating the same request. A rate limit may require backoff. A provider-specific error contract must decide the category.
2. Is the operation idempotent?
An idempotent operation can be repeated with the same intent without creating an unintended second effect. The downstream API may support an idempotency key, a natural business key, or a read-before-write reconciliation flow. Do not assume that a tool named create_* is safe to retry.
3. What should happen when the outcome is unknown?
The harness should record an unknown_outcome state, query the downstream system if possible, and escalate when it cannot establish the result. That is safer than letting the model improvise.
A practical operation record might include a stable operation ID, idempotency key, target resource, request hash, current status, provider response ID, and reconciliation result. The exact fields depend on the provider and workload, but the principle is stable: the agent must be able to distinguish “not attempted” from “attempted but unknown.”
| Situation | Unsafe reaction | Harness response |
|---|---|---|
| Validation error | Retry the same arguments | Stop, repair only with new validated input |
| Transient read failure | Retry without a cap | Bounded backoff, budget, and cancellation |
| Timeout after mutation | Assume failure and repeat | Mark unknown, reconcile, then decide |
| Permission denial | Ask the model to try harder | Stop or request an approved capability |
Retry logic without side-effect semantics is not reliability. It is duplication with better marketing.
State, memory, and the difference between resumable and reversible
An agent may need information at three different scopes:
- Turn context: what is needed for the current model call.
- Thread or session state: what is needed to continue a task.
- Long-term memory: information that should survive the task and be reused later.
These scopes should not be treated as one giant conversation history. LangGraph distinguishes thread-scoped checkpoints from cross-thread stores, and its documentation notes that in-memory persistence disappears on restart. Vertex Frontier’s coverage of agent memory architecture goes deeper into provenance, authority, freshness, and read/write policy.
The harness should define who owns each state, how sensitive it is, how long it lives, how it is corrected, and whether a later action may rely on it. A vector index is not automatically authoritative memory. Retrieved text is evidence to evaluate, not an instruction to obey.
There is another distinction worth making:
A checkpoint can restore logical workflow state. It cannot automatically undo an external side effect.
If the agent saved “step 3 complete” before a payment was committed, replaying from the checkpoint may repeat or skip the payment. Recovery needs side-effect records and reconciliation, not just serialized conversation history.
Long context creates a related failure mode. More tokens do not guarantee better decisions. Context can become stale, noisy, or internally contradictory. See Vertex Frontier’s analysis of context rot in LLMs and its long-context benchmarking guide for the measurement side of the problem.
Durable state needs ownership and recovery semantics. “We saved the checkpoint” is not the same as “we can safely resume.”
Sandboxing: capability is useful, but capability also expands the blast radius
A sandbox can restrict filesystem access, processes, network routes, credentials, or runtime privileges. It can make coding, browsing, data transformation, and testing safer. But “the agent runs in a container” is not a complete security statement.

Isolation depends on configuration and environment:
- Which user does the process run as?
- Can it reach the public internet or internal services?
- Are cloud credentials mounted into the environment?
- Is the filesystem disposable or shared?
- Can the process create child processes?
- What happens if the agent receives malicious instructions from a file or webpage?
- Can a tool escape the intended boundary through a privileged helper?
Google ADK documents sandboxed execution and network controls, while NVIDIA OpenShell’s overview describes policy-based isolation. OpenHands and other coding-agent environments show why stateful computer access can be valuable for complex tasks. These sources describe particular designs, not a universal guarantee.
The contrarian point is that a more capable computer is not automatically a better agent. It may complete more tasks while increasing the number of things that can go wrong. Capability and blast radius rise together unless permissions, network paths, credentials, and review gates are designed with equal care.
For high-impact actions, separate observation from mutation. Let the agent inspect data, prepare a plan, and produce a diff before it can commit a change. Require explicit approval for actions that are difficult to reverse.
A sandbox is a boundary, not a magic word. Describe what it isolates and test the boundary under hostile conditions.
Threat modeling the harness: assume the tool result can be hostile
An agent does not operate in a single trusted conversation. User text, retrieved documents, web pages, tool results, memory, model output, and external services can all cross trust boundaries. Treating every returned string as an instruction is one of the fastest ways to turn a useful tool into an unintended control channel.
A practical threat model asks four questions for every data flow:
AI Agent Threat Matrix & Harness Controls
Scroll horizontallyKey vulnerabilities, signatures, and mitigations for secure agentic systems.
| Threat | What it looks like | Harness control |
|---|---|---|
| Prompt injection | A webpage or document tells the agent to ignore its task and reveal data | Keep retrieved content as untrusted data; enforce policy outside the model and restrict tool scope |
| Tool or result poisoning | A tool result contains instructions, altered fields, or a misleading success message | Validate schemas, bind results to the requested resource, and verify postconditions |
| Excessive agency | A read-only task unexpectedly gains access to delete, send, publish, or deploy tools | Use task-scoped capability sets and separate read, prepare, and commit operations |
| Credential or data exfiltration | Secrets or sensitive context are passed to an untrusted tool or external host | Use short-lived identity, destination controls, redaction, and explicit data-flow review |
| Cross-tenant access | A valid tool call targets another customer’s record | Authorize the target server-side and derive tenant identity from trusted context, not model arguments |
This is not a claim that one filter eliminates these risks. It is a design habit: identify the untrusted input, name the privileged action it could influence, and place an enforceable control between them. The OWASP AI Agent Security Cheat Sheet is a useful reference for excessive agency, tool misuse, prompt injection, and least-privilege concerns.
For an agent connected through MCP or another tool protocol, treat the protocol as a transport and capability boundary, not as proof that a server is trusted. Review which tools are exposed, what data leaves the process, how authorization is delegated, and whether the server can reach additional systems. Vertex Frontier’s guide to MCP security, permissions, and OAuth hardening provides a focused follow-up for that review.
The harness should assume that text can be adversarial even when it arrived through a tool that appears useful.
Observability: record the run, not a fictional story about reasoning
An agent can return HTTP 200 and still fail the task.
The model may choose the wrong tool, use a stale record, repeat a call, exceed a budget, or report success without checking the result. This is why agent observability must cover the trajectory as well as the final response.
The OpenAI tracing documentation covers runs, turns, agents, generations, tools, guardrails, handoffs, and custom events. Google’s agent observability guidance similarly focuses on actions, tool choices, traces, quality, safety, and reliability.
A useful structured event can include:
- run ID and parent span;
- model and application version;
- context or prompt-template version;
- tool name and version;
- redacted arguments and target resource;
- agent identity and approval state;
- timestamps, latency, retry count, and error class;
- side-effect status and provider response ID;
- final outcome and evaluator result.
Do not confuse three different records:
- Application trace: explains what the runtime did.
- Audit record: provides a controlled, durable record for accountability.
- Hidden model reasoning: internal model content that should not be treated as a required audit artifact.
Good observability does not require exposing private chain-of-thought. It requires recording the decisions and actions that the application can verify.
Vertex Frontier’s MLflow observability guide for multi-agent systems is a useful internal destination for trace, span, cost, privacy, and tool-call detail.
Debuggable agents leave an evidence trail. The final answer is only one event in that trail.
Production incident runbook: what to do when an agent goes wrong
Reliability is not only the ability to prevent failures. It is the ability to investigate and contain them without asking the model to guess what happened. Keep an operator-facing runbook close to the harness design.
When a customer reports a suspicious action or a run stops unexpectedly:
- Freeze further mutation if necessary. Disable the affected capability or move it to approval-required mode; do not begin by replaying the run.
- Locate the run and operation IDs. Find the application trace, tool invocation, approval event, provider response identifier, and side-effect record.
- Classify the terminal state. Distinguish
blocked,failed_before_effect,committed,partially_completed, andunknown_outcome. - Reconcile with the downstream system. Query the source of truth using the stable operation or business identifier. A local timeout is not evidence that no mutation occurred.
- Contain the blast radius. Revoke or rotate exposed credentials, narrow the tool catalog, isolate affected tenants, or stop the relevant workflow if the issue crosses a trust boundary.
- Replay safely. Reproduce in a read-only or sandboxed environment first. If a mutation must be retried, use the provider’s idempotency or reconciliation mechanism.
- Record the new evaluation case. Add the failure trajectory, relevant tool output, policy decision, and expected recovery behavior to regression and adversarial tests.
The exact response depends on the domain. A support refund, source-code change, deployment, and medical workflow have different owners and escalation paths. The reusable part is the separation between containment, reconciliation, recovery, and learning.
A harness is not production-ready until an operator can answer “what happened?” and “what should we do next?” without relying on the model’s memory.
Evaluation: judge the path, not only the answer
Traditional software tests often check input and output. Agents need that layer, but it is not enough.

Suppose an agent reaches the correct answer after exposing a secret to an untrusted tool, attempting three unauthorized actions, and getting lucky on the fourth call. A final-answer evaluator may mark it as successful. A production evaluator should not.
Google ADK’s evaluation documentation recommends evaluating tool choice, strategy, trajectory, efficiency, final response, and session state. OpenAI’s evaluation best practices emphasize task-specific datasets and graders, while LangChain’s agent evaluation resources describe evaluation across runs, traces, and threads.
A practical evaluation stack has at least five layers:
- Unit tests: validators, permissions, parsers, and policy decisions.
- Integration tests: model-to-tool contracts, state transitions, and cancellation.
- Replay tests: restart, resume, duplicate delivery, and unknown outcomes.
- Adversarial tests: prompt injection, malicious tool output, data exfiltration, and privilege escalation.
- End-to-end tests: task completion, user-visible quality, cost, latency, and human override.
Measure more than answer quality. Track unsafe-action rate, unnecessary tool calls, retry rate, recovery rate, latency, cost, escalation rate, and the percentage of runs that end in an unknown state. Segment results by task class because a support agent and a deployment agent do not carry the same risk.
This is also where the idea of harness lift becomes useful. If a new model improves the result, ask whether it improved the model or whether the harness improved the context, tools, checks, and recovery path. Vertex Frontier’s discussion of recursive language models and harness lift provides a relevant benchmark caveat: model scores do not automatically measure production reliability.
A release gate for agents
Before allowing an agent to act on real data, require evidence for each high-impact capability:
- the tool has a defined owner and permission model;
- invalid arguments are rejected before execution;
- retries have a documented side-effect policy;
- unknown outcomes are represented and reconciled;
- traces are redacted and queryable;
- recovery has been tested after process failure;
- adversarial tool outputs have been tested;
- human approval is required where reversibility is low;
- the evaluation set includes realistic failures, not only happy paths.
A passing demo is a starting point. A release gate is a claim backed by tests and evidence.
Evaluate what the agent did, what it changed, what it cost, and how it recovered, not only what it said.
Before and after: how a harness changes the product
Consider a customer-support agent that can search an order, issue a refund, and send an email.
Before the harness: the model receives the conversation and tool descriptions. It chooses a tool, the application executes it, and the model writes a reply. There is no stable operation ID, no explicit refund approval threshold, no postcondition check, and no reliable way to tell whether a timeout happened before or after the refund.
After the harness: the runtime verifies the customer identity, limits tool exposure, validates the order and amount, attaches an idempotency key, pauses for approval above a threshold, records the tool call, reconciles unknown outcomes, verifies the refund status, and gives the model only the sanitized result needed to explain the outcome.
The second system may feel less autonomous. It is more useful because it is accountable.
Model proposes a refund → application executes → model reports success.
Risk: duplicate or unauthorized mutation with weak evidence.
Validate identity and policy → approve exact action → execute with idempotency → reconcile → verify postcondition → report.
Benefit: controlled action with a recoverable record.
Common mistakes in agent harness engineering

Mistake 1: Treating a system prompt as a security boundary
Prompts are changeable inputs. They can be ignored, contradicted, or bypassed by a tool result. Put authorization and limits in enforceable code or policy controls.
Mistake 2: Retrying every failure
A retry policy needs an error classification and side-effect contract. Invalid input, permission denial, and unknown mutation outcomes are not the same failure.
Mistake 3: Logging only the final answer
The final answer cannot explain why a tool was called, which arguments were used, whether approval happened, or whether the state was already mutated. Capture the trajectory with redaction.
Mistake 4: Treating checkpoints as rollback
A checkpoint can restart a workflow. It does not reverse a payment, email, deployment, or database mutation. Pair checkpoints with operation records and reconciliation.
Mistake 5: Giving every tool to every agent
Capability should be scoped to the task, user, tenant, and stage of execution. Tool discovery is not permission.
Mistake 6: Measuring only task completion
A system that completes a task by violating policy is not reliable. Add unsafe-action, unnecessary-call, recovery, and human-override measures.
Mistake 7: Calling a container a secure sandbox
Test credentials, network routes, filesystem access, process privileges, and hostile inputs. State the boundary precisely.
Mistake 8: Choosing an agent when a workflow is enough
Dynamic autonomy adds uncertainty. If the path is known, a controlled workflow may be easier to test, operate, and explain. Use agency where it creates real value, not because the architecture diagram looks more advanced.
A practical maturity model
Teams can use the following model to decide what they actually have:
- Level 0 — Model call: prompt in, answer out.
- Level 1 — Tool loop: the model can call tools, but limits and permissions are informal.
- Level 2 — Controlled execution: schemas, budgets, validation, and explicit terminal states exist.
- Level 3 — Recoverable operation: state, idempotency, reconciliation, and restart behavior are tested.
- Level 4 — Evidenced production system: traces, adversarial evaluations, release gates, ownership, and incident procedures are in place.
This is not an industry certification. It is a diagnostic model. A read-only research assistant may be valuable at Level 2. A deployment or financial agent should demand stronger evidence before it receives mutation rights.
A production-readiness matrix you can actually use
The maturity levels above are useful for conversation, but teams still need a concrete release decision. The matrix below turns the idea into a review artifact. It is a decision checklist, not a benchmark and not a claim that any particular framework satisfies the controls automatically.
| Control area | Evidence to require before mutation access | Release question |
|---|---|---|
| Capability | A versioned tool catalog with owners, schemas, allowed targets, and explicit side-effect labels | Can the team name exactly what this agent is allowed to change? |
| Authorization | A test showing that denied users, tenants, paths, and amounts are rejected outside the prompt | Would the action remain blocked if the model ignored its instructions? |
| Side effects | Idempotency or reconciliation behavior for timeouts, duplicate delivery, and partial completion | What happens when the client cannot tell whether the write succeeded? |
| Recovery | Replay or restart tests with explicit `unknown_outcome`, `blocked`, and `needs_review` states | Can an operator resume safely without repeating a committed action? |
| Evidence | Redacted traces linking tool calls, approvals, identities, errors, and final outcomes | Can the team reconstruct what happened after a complaint? |
| Evaluation | Adversarial, trajectory, recovery, and end-to-end cases for the actual workload | Does the test set include unsafe paths that still produce a correct final answer? |
If a row has no owner or evidence, the missing item is not “more autonomy.” It is a release blocker for that capability. Start with read-only access, collect the evidence, and grant mutation rights one tool or operation class at a time.
Use the production-grade agent harness walkthrough on Vertex Frontier when you need an implementation-oriented architecture. Use this article’s model to decide which controls and evidence that architecture must demonstrate.
Choosing between a workflow and an agent
The question is not “Which framework has the most features?” It is “Where does dynamic decision-making create enough value to justify the operational cost?”
Choose a workflow when:
- the path is predictable;
- each step has a clear contract;
- deterministic routing is easy to maintain;
- failures must be tightly bounded.
Choose an agent when:
- the environment changes during execution;
- the next useful action cannot be enumerated in advance;
- the agent must select among tools or strategies;
- the value of flexibility exceeds the cost of evaluation and control.
A hybrid is often the sensible choice: deterministic policy and commit stages around a model-driven planning or information-gathering stage.
For framework-specific context, Vertex Frontier’s Google Agent Development Kit production guide is best treated as an example of how one framework instantiates these concepts, not as proof that a framework removes the engineering work.
What LLM vendors give you – and what they do not
Model and agent vendors provide increasingly useful building blocks: tool schemas, runners, guardrails, tracing, sessions, evaluation APIs, and reference patterns. Those features reduce implementation effort.
They do not know your business invariants, tenant boundaries, downstream idempotency behavior, approval policy, incident response process, or acceptable risk. They cannot decide whether an unknown payment outcome should be reconciled automatically or escalated to a human. They cannot make a generic sandbox safe for your credential model.
That work remains product-specific.
The most durable architecture is therefore not “vendor feature versus custom code.” It is a clear ownership map:
- the provider owns its API contract and documented runtime behavior;
- the framework owns the abstractions it actually implements;
- your harness owns business policy, identity, side-effect safety, recovery, evidence, and release decisions.
Final takeaway
Agent harness engineering is the discipline of making model-driven behavior accountable to the real world.
It turns an open-ended model call into a bounded operation with permissions, state, recovery, evidence, and a clear stopping rule. It treats a timeout after a mutation as an unknown outcome, not a convenient failure. It treats a prompt as guidance, not authorization. It evaluates the path, not only the final sentence.
The practical test is simple: if your model can call a tool that changes something important, ask what happens when the tool is slow, wrong, duplicated, compromised, or only partly successful.
If the answer is “the model will figure it out,” you have a prototype.
If the answer is backed by policy, state, reconciliation, traces, tests, and an owner, you are doing harness engineering.
FAQs
Is agent harness engineering the same as an AI agent framework?
No. A framework may provide runners, tools, state, guardrails, or tracing. Harness engineering is the broader discipline of designing and operating the complete control system, including business permissions, side-effect semantics, recovery, evaluation, and ownership. A framework is one possible implementation component.
What is the difference between an LLM API call and a production agent?
An API call produces an output. A production agent runs inside a controlled loop that can select or receive tool actions, validate them, execute permitted operations, maintain state, handle failure, observe the trajectory, and stop or escalate under explicit rules.
Why is idempotency important for AI agents?
Agents operate across networks and may retry after timeouts. If a mutation succeeded but the response was lost, a blind retry can create a duplicate side effect. An idempotency key or reconciliation process helps the system determine whether the original operation already happened.
Do guardrails make an AI agent safe?
Guardrails reduce specific risks, but they are not a universal safety guarantee. Effective protection also requires authorization, least privilege, input and output validation, tool controls, isolation, postcondition checks, monitoring, adversarial evaluation, and a response plan for failures.
What should agent observability capture?
Capture the run and action trajectory: run ID, model and application versions, tool names, redacted arguments, identity, approvals, timestamps, latency, retries, errors, side-effect status, provider response identifiers, and final outcome. Keep application traces separate from durable audit records and hidden model reasoning.
Should every AI agent run inside a sandbox?
The answer depends on the capabilities and data involved. Sandboxing is especially important when an agent can execute code, access files, browse the web, or reach external systems. The boundary must be tested and described precisely; a container alone does not prove secure isolation.
How do you evaluate an agent beyond its final answer?
Evaluate tool choice, arguments, approvals, state transitions, unnecessary calls, retries, policy violations, recovery, latency, cost, and final outcome. Use unit, integration, replay, adversarial, and end-to-end tests. A correct final answer does not excuse an unsafe path.
Was this article helpful?










[…] a name now a harness, and it’s the part almost nobody is actually showing you how to build. Plenty of pages define the term. Almost none of them ship […]
[…] This article develops the conceptual model. For an updated treatment that focuses specifically on the production discipline, the runtime controls, idempotency, side-effect safety, threat modeling, and the maturity model teams use to decide when an agent is ready for mutation rights, read our companion piece on what agent harness engineering means in production. […]
[…] For a working definition of agent harness engineering as a production discipline, one that separates the model, the prompt, the tool layer, the runner, the runtime, and the harness as the engineering layer that makes them work together, see our primer on what harness engineering actually covers and what it does not. […]
[…] discipline is one layer of a broader practice, agent harness engineering, which covers the control loop, guardrails, tool contracts, retries, state, sandboxing, and […]
[…] qualification: ADK gives you the building blocks for a production system; it does not make the system production-ready by itself. The difficult work appears at the boundaries, between a local development UI and a deployed […]
[…] that gap is the job of agent harness engineering: the production discipline that wraps every tool call, every retry, every approval, and every side […]
[…] “you build it” layer is exactly what agent harness engineering names and structures: the runtime, control loop, guardrails, tool contracts, retries, state, […]