Your agent worked perfectly in the demo. It answered questions, called the right tools, sounded confident. Then it went live, and within a week someone found a conversation where it called the same refund tool four times in a row, a support ticket where it “remembered” a promise it never made, and a Slack message from finance asking why last week’s model bill looked nothing like the estimate. None of that is a sign the model got worse. It’s a sign nothing was watching it.
That gap between an agent that works in a demo and one that survives production is almost never about the model. It’s about everything wrapped around the model: the part that remembers what happened last Tuesday, the part that notices a tool call is about to run for the fortieth time, the part that can prove, after the fact, exactly why the agent did what it did. That wrapper has a name now a harness, and it’s the part almost nobody is actually showing you how to build. Plenty of pages define the term. Almost none of them ship one.
This article does. What follows is a complete, code-first build: a typed state schema, memory tiered like an operating system’s, retries and circuit breakers with real failure modes, evals wired into your deploy pipeline, and full tracing, all built around one running example, a customer-support agent, so every design decision has somewhere concrete to land. No toy REPL loop, no framework doing the hard parts off-screen. Just the harness, built once, in the open.
What you will get from this guide
Click any topic to expand or collapseHarness vs. Model & Framework
A harness, not the model, decides whether an agent survives contact with production. We define what a harness owns versus what a framework or the model owns, and why that split matters.
State First Design & Pydantic Schema
State comes first. We design a typed, versioned Pydantic schema that every other subsystem — retries, evals, tracing — reads and writes, instead of letting state leak through untyped dicts.
Three-Tiered Memory Architecture
Memory gets three tiers — hot, warm, cold — each with its own store, token budget, and eviction rule, plus the promotion/demotion code that moves data between them.
Reliability Chain & Safety Guardrails
Reliability is a chain, not a single trick: retry for transient errors, a circuit breaker for a sustained provider outage, and a budget-and-kill-switch watchdog for runaway loops that neither retries nor breakers catch.
Evals as a Deploy Gate
Evals get wired as a deploy gate — three layers of testing, from deterministic invariants to weekly production sampling — with production failures feeding back into the test set.
Full-Loop Instrumentation & Tracing
The whole loop gets instrumented with OpenTelemetry’s GenAI semantic conventions, so a support ticket’s trace connects the model call, the tool call, and the retrieval step under one ID.
Grounded Customer-Support Case Study
Every design choice is grounded in a customer-support agent, because side-effecting tools (refunds, ticket updates) and measurable outcomes (resolution, deflection, escalation) make the trade-offs concrete instead of abstract.
What this article is, and isn’t
This is a design walkthrough of one harness architecture, built from official documentation, vendor engineering guidance, and patterns practitioners repeatedly describe in production. It is not a benchmark report — we haven’t run this exact system at scale and won’t pretend otherwise. Where a number comes from a vendor’s own case study, we say so and link it; where a claim comes from a community thread, we treat it as one practitioner’s account, not proof.
The Harness Decides Whether Your Agent Survives
Practitioners who move agents from a demo to production keep hitting the same wall, and it’s rarely the model. State gets passed between steps with no schema, so a typo in a dictionary key fails silently three calls later. A tool call fails, the agent retries, gets a slightly different error, and tries again, sometimes for hours, sometimes until a billing alert fires. Memory drifts: the agent “remembers” something that was never true, or forgets something it was told an hour ago, and nobody can point to where it went wrong.

None of that is a model problem. A demo agent is just a loop: call the model, run a tool, feed the result back, repeat. A production agent is everything wrapped around that loop, the part that persists state when the process restarts, the part that stops a call before it becomes the fortieth retry, the part that proves the system got better (or worse) after last week’s prompt change. That wrapper is the harness. Microsoft’s Agent Framework documentation describes it as the runtime scaffolding that drives model and tool calls, manages conversation state, and applies approval policies; the model is rented, but the harness is what you actually own and maintain.
Customer support is a deliberately hard test case for a harness, and that’s the point. The agent runs multi-turn conversations, calls tools with real side effects (issuing a refund, updating a ticket), and has to know when to stop and hand off to a person. Anthropic’s engineering guidance on building effective agents names customer support as one of the clearest production use cases for agents, precisely because the outcomes, resolved, escalated, refunded, are things you can actually measure. If a harness design holds up here, it holds up almost anywhere.
This article walks through six build steps in the order you should actually do them: state, memory, reliability, evals, observability, and shipping. Each step includes working Python, a table or two, and the specific failure modes the step exists to prevent. By the end you’ll have a framework-agnostic core you can drop LangGraph, the OpenAI Agents SDK, or nothing at all underneath.
What Is an Agent Harness (and What It Isn’t)
The clearest one-line distinction practitioners and vendors converge on is this: a framework is for building an agent; a harness is for running one. A framework gives you abstractions for wiring a model to tools and prompts. A harness is the layer that actually executes those wired-up pieces in production, it enforces budgets, persists state between turns, and decides what happens when a tool call fails at 2 a.m.

Martin Fowler’s essay on harness engineering frames the same idea through control theory: a harness is a feedback loop, not a static prompt. It watches what the model does (feedforward), compares that against what should happen, and adjusts (feedback), closer to a thermostat than a script. That framing matters because it means the harness’s job never finishes at deploy time; it keeps regulating behavior for as long as the agent runs.
Harness vs. agent vs. framework, concretely:
- The agent is the model plus its instructions and available tools, the thing that decides what to do next.
- The framework (LangGraph, the OpenAI Agents SDK, CrewAI) gives you pre-built plumbing for wiring the agent together, graphs, handoffs, state containers.
- The harness is the runtime that actually executes runs: it owns persistence, retry policy, budget enforcement, evaluation, and tracing, whether or not a framework sits underneath it.
You can build a harness with a framework’s plumbing or without it, harness engineering treats the model as one component inside a control system, and that’s the mental model this article builds on top of, one subsystem at a time.
The Architecture at a Glance
Read the rest of this article as building seven subsystems, in this order, because each one depends on the one before it:
System Architecture Breakdown
Click any component to expand or collapse details1. Intake
The entry point that receives a customer message and starts or resumes a run.
2. State store
A versioned, typed schema in managed Postgres; everything else reads and writes here.
3. Agent loop
A finite-state machine (intake → retrieve → draft → verify → respond/escalate) that advances the state.
4. Memory tiers
Hot (Redis), warm (Postgres + pgvector), and cold (object storage) for context of different ages.
5. Resilience layer
Retries, a circuit breaker, and budget/kill-switch guards wrapped around every model and tool call.
6. Evals
A CI-gated test suite plus weekly production sampling, feeding failures back into the test set.
7. Observability
OpenTelemetry spans and metrics tying one trace_id to every model call, tool call, and retrieval in a run.
State comes first in the build order because every other subsystem reads or writes it. The retry layer needs somewhere to record attempt counts. The evaluator needs a versioned transcript to replay. The tracer needs a stable run ID to hang spans off. Skip the schema and you're improvising all three later, usually under production pressure, which is exactly what practitioners report happening when state is untyped and shared ad hoc between components.
Scope and assumptions
Click any point to expand or collapse details1. Conversation Thread Scope
One agent run corresponds to one customer conversation thread; the compaction thresholds in Code Block B (12 turns / 4,000 tokens) are starting points to tune against your own conversation-length distribution, not fixed numbers.
2. Side-Effecting Tools
At least one tool call has a real side effect — a refund, a ticket update. If every tool in your agent is read-only, the idempotency-key and tool-risk-tiering work in Steps 3 and 6 matters less; you can simplify.
3. Provider Account Control
You control your own model-provider account and can read its actual rate-limit and spend-cap documentation — the retry and breaker thresholds in Step 3 have to be tuned per provider, never copy-pasted blind.
4. Infrastructure Ownership
You're able to operate Postgres and Redis, or managed equivalents. This harness assumes you own that infrastructure; it isn't scoped for a fully serverless, zero-ops target.
5. Latency Expectations
This isn't designed for sub-200ms voice or real-time interaction. The retrieve-draft-verify sequence adds latency a support chat can absorb but a voice agent usually can't.
Repository at a Glance: One File per Decision
Before we build the subsystems one by one, here is where each one lives in the repository. Keeping one file per decision is deliberate: when a retry policy misbehaves at 2 a.m., you want the answer sitting in resilience.py, not smeared across a god-class called agent.py.
The tree below is the map for the rest of this walkthrough, every code block that follows lands in exactly one of these files, and every file maps back to one of the seven subsystems above.
support-harness/
├── pyproject.toml # Python 3.12 — pydantic, tenacity, opentelemetry-api, psycopg, redis, pytest
├── docker-compose.yml # Local stack: Postgres (+pgvector), Redis, OTel collector
├── src/harness/
│ ├── state.py # Subsystem 2 — AgentState, Event, ControlFlags, Checkpoint (Code A)
│ ├── state_store.py # Versioned Postgres persistence: append-only events, upcasting on read
│ ├── memory.py # Subsystem 4 — hot/warm/cold tiers, summarize_and_promote() (Code B)
│ ├── llm_client.py # The only file that knows the model provider — swap it, nothing else changes
│ ├── resilience.py # Subsystem 5 — @retry policy (Code C), CircuitBreaker (Code D), BudgetGuard
│ ├── tools.py # lookup_order, issue_refund, create_ticket — each with a risk tier + idempotency key
│ ├── escalation.py # Human handoff as a designed state: context bundle, idempotent handoff
│ ├── loop.py # Subsystem 3 — the FSM turn loop: state in → model → tools → state out
│ └── observability.py # Subsystem 7 — OTel GenAI span wrapper (Code E), cost receipts per run
└── evals/
├── conftest.py # Golden-trajectory fixtures: runtime profile + frozen production replays
├── test_invariants.py # Layer 1 — deterministic: allowed tools, valid JSON, token/latency ceilings
├── test_scenarios.py # Layer 2 — support scenarios incl. refusal and escalation correctness
└── test_regression.py # Layer 3 — frozen production failures; CI blocks the merge on regression
Three design choices are worth naming before we write any code.
- First, llm_client.py is the only module that knows which provider you use, swap OpenAI for Anthropic or a local model and nothing else changes, which keeps the core framework-agnostic in substance rather than just in the README.
- Second, evals/contains ordinary pytest tests instead of notebooks; that is what lets CI block a merge when resolution quality regresses, and it is the whole point of Step 4.
- Third, escalation.py sits beside tools.py rather than inside it, because human handoff is a designed state in the schema (Step 1), not an error branch buried in tool code. If you later adopt LangGraph, you replace loop.py and state_store.py with its checkpointer and keep the remaining seams exactly where they are, the same trade we examine in "Build vs Adopt".
Step 1: Define the State Schema First

Why schema-first
When an agent state is just a shared dictionary passed between functions, failures go quiet instead of loud. A key gets renamed in one place and not another; nothing throws an exception, the run just silently produces a wrong answer three steps later. One practitioner's summary of debugging LangGraph agents in production put it plainly: untyped state passed between nodes with no schema is the first thing that breaks, and it breaks silently. A typed schema turns that class of bug into a validation error at the boundary, where you can actually see it.
There's a second reason to do this first: retries, evals, and tracing all need to read state to do their job. The retry layer needs a place to store attempt counts and idempotency keys. The evaluator needs a stable, versioned transcript to replay. The tracer needs one run ID to attach every span to. Define the schema before any of those subsystems, and they all have somewhere consistent to look at.
What state actually contains
A harness's state needs, at minimum: an append-only transcript of everything that happened in the run, a ledger of which tools were called and with what idempotency keys, control flags (how many steps taken, how many tokens spent, whether the run has hit its budget), an escalation state for the human-handoff path, and a schema version so you can evolve the shape later without breaking old runs.
Each of those choices maps to a documented pattern, not a guess. Anthropic's guidance on effective agents recommends giving agentic loops explicit stopping conditions, a maximum number of iterations, precisely so control doesn't depend on the model deciding to stop.
Stripe's idempotent-requests documentation is the direct model for the idempotency-key ledger: a side-effecting call (a refund, in our case) stores the key alongside the result of the first attempt, so a retried call returns the original result instead of firing twice. And the append-only transcript follows the event-sourcing pattern documented in Microsoft's Azure Architecture Center: store the full sequence of events, derive current state by replaying them, and get an audit trail for free.
What this code assumes
| Library | Version / API used here | Role in the harness |
|---|---|---|
| Python | 3.12-style syntax (built-in generics) | Core language for every snippet |
| Pydantic | v2 (ConfigDict, model_dump_json) | State schema validation and serialization |
| tenacity | current release (wait_exponential_jitter) | Retry with backoff and jitter |
| redis-py | compatible with Redis 7.x EXPIRE | Hot-tier client |
| psycopg + pgvector | pgvector with HNSW index support | Warm-tier store and retrieval |
| opentelemetry-api / sdk | GenAI semconv, Development stability | Tracing and metrics |
The state schema code
Python — state schema (state.py):
from datetime import datetime
from uuid import UUID, uuid4
from pydantic import BaseModel, ConfigDict, Field
class Event(BaseModel):
model_config = ConfigDict(frozen=True)
kind: str # user_msg | tool_call | tool_result | model_msg
payload: dict
at: datetime = Field(default_factory=datetime.utcnow)
class ControlFlags(BaseModel):
step_count: int = 0
max_steps: int = 12
tokens_spent: int = 0
token_budget: int = 20_000
escalation_state: str = "none" # none|pending|handed_off
class AgentState(BaseModel):
model_config = ConfigDict(extra="forbid")
schema_version: int = 2
run_id: UUID = Field(default_factory=uuid4)
transcript: list[Event] = Field(default_factory=list)
control: ControlFlags = Field(default_factory=ControlFlags)
idempotency_keys: dict[str, str] = Field(default_factory=dict)
def dump(self) -> str:
return self.model_dump_json()Each Event is frozen, once written, it doesn't change, which is what makes replay and audit trustworthy. extra="forbid" on the top-level model means a stray field anywhere in the pipeline raises immediately instead of silently riding along. This is the trade Pydantic v2's model validation makes for you automatically: it validates that the parsed output actually matches the declared types, coercing where it safely can and raising where it can't.
Versioning without breaking old runs
Never mutate history directly, if a fact turns out to be wrong, append a compensating event rather than editing the old one, the same discipline event-sourced systems use to keep an audit trail honest.
When the schema itself changes, three techniques from the event-sourcing pattern keep old runs readable: tolerant deserialization (ignore fields you don't recognize, default the ones you're missing), an explicit version field so you know which shape you're looking at, and upcasting, small, chainable functions that convert an old event shape into the current one on read. None of this is exotic; it's the same versioning discipline any event-sourced system needs, just applied to agent runs instead of order histories.
Where state lives
Put the authoritative copy in managed Postgres, never on a process's local or ephemeral disk. A community thread that's become a small cautionary tale describes exactly this failure: a LangGraph SQLite checkpointer deployed on Vercel lost its entire memory on every redeploy, because the SQLite file lived on ephemeral storage that doesn't survive a new deployment. The fix reported by the thread was the same one this article recommends by default, move the checkpointer to managed Postgres and keep a versioned shadow snapshot for rollback.
LangGraph's persistence documentation reaches a compatible conclusion from a different direction: it splits state into a checkpointer (thread-scoped, short-term, used for conversation continuity and fault tolerance) and a store (cross-thread, long-term memory), and recommends PostgresSaver over the in-memory saver specifically because in-memory checkpoints vanish on restart. Whether or not you adopt LangGraph, that split, durable per-thread state versus durable cross-thread memory, is worth keeping, because it's the same distinction our hot/warm/cold memory tiers make in the next step.
A framework note, for honesty's sake: if you do use LangGraph, its documented Graph API default schema type is TypedDict, not Pydantic, the docs note that Pydantic validation is measurably slower than a TypedDict or dataclass, and the higher-level create_agent factory doesn't support Pydantic state at all. That's a real trade-off. We use Pydantic here anyway, because the validation guarantees, catching a malformed field at the boundary instead of three steps downstream, are worth the extra microseconds for a support agent that touches refunds and tickets.
Step 2: Memory in Three Tiers: Hot, Warm, Cold

The problem context windows create
Every context window is finite, and quality doesn't hold steady as it fills up. Chroma's research on context rot tested across 18 models, found that model performance consistently degrades as input length grows, and that the effect gets worse when the added context contains distractors rather than clean signal.
Anthropic's guidance on effective context engineering describes the same thing as an "attention budget": every token you add spends down a limited resource, so the practical fix isn't a bigger window, it's curating what goes in.
That's also why stuffing more history into the prompt degrades quality (context rot) instead of simply making the agent smarter, more isn't better once you're past the point where the model can actually attend to all of it.
The harness answer to that constraint is tiering: not everything the agent has ever seen needs to sit in the live prompt. Anthropic's context-engineering guidance describes compacting older turns into summaries and loading detail "just-in-time" by reference rather than pre-loading it, exactly the pattern our tiers implement in code.
Tier definitions and the budget table
| Tier | Store | Holds | Budget | Eviction |
|---|---|---|---|---|
| Hot | Redis + in-process | Current thread scratch, recent tool results | Per-run token cap (~2–4k words) | TTL + allkeys-lru |
| Warm | Postgres + pgvector | Per-thread summaries, resolved entities, knowledge base | Top-k retrieval (ef_search default 40) | None — budgeted by k, not evicted |
| Cold | Object storage / warehouse | Full transcripts, audit trail, eval replay datasets | Loaded on demand by reference | Retention policy, not cache eviction |
Redis's documentation on key eviction frames allkeys-lru as the sensible default for most workloads under the 80/20 rule, which is why it's the hot-tier default here; Redis's EXPIRE command handles the TTL side with per-key precision. pgvector's README documents HNSW as the practical warm-tier default, better recall-per-speed trade-off than IVFFlat, and it can be built before any data exists in the table, which matters for a fresh deployment.
Promotion and demotion
Python — memory tiering (memory.py):
def summarize_and_promote(state: AgentState, hot: RedisClient, warm: PgVectorStore,
turn_threshold: int = 12, token_threshold: int = 4_000) -> AgentState:
recent = state.transcript[-turn_threshold:]
used_tokens = sum(len(e.payload.get("text", "")) // 4 for e in recent)
if len(state.transcript) < turn_threshold and used_tokens < token_threshold:
return state # hot tier still under budget, nothing to compact
summary = llm_summarize(state.transcript[:-turn_threshold]) # structured, entity-tagged
warm.upsert(thread_id=state.run_id, summary=summary, at=datetime.utcnow())
compacted = state.model_copy(update={
"transcript": [Event(kind="summary_ref", payload={"warm_id": str(state.run_id)})] + recent
})
hot.set(f"thread:{state.run_id}", compacted.dump(), ex=3_600)
return compactedThis mirrors the shape of Anthropic's context-editing beta, which triggers a server-side summarization pass once input tokens cross a configurable threshold (100,000 by default, as documented as of September 2026) and keeps a limited number of recent tool uses in full. You don't need the beta feature to get the same effect, the function above does it explicitly, in your own infrastructure, which also means you control exactly what gets kept, summarized, or dropped.
One of the most detailed first-hand accounts in the community data — from a team that had built memory into four separate agent systems — put it bluntly: a vector store alone gets you keyword search with extra steps and worse debugging. What reportedly worked instead was entity resolution (even naive name-normalization gets most of the way there), tagging facts with when they were true (a decision from three months ago might already be reversed), and surfacing contradictions to a human instead of trying to auto-resolve them.
Build the warm tier as structured summaries and resolved entities, not a raw transcript dump, the knowledge base the warm tier retrieves from is worth building carefully, because retrieval quality is bounded by what you feed the index, not by which vector database you picked.
The cold tier does double duty. It's the replay source for the evals pipeline in Step 4, a versioned archive you can run failed production runs back against after a fix, and it's the compliance archive. That second job creates a real tension: an append-only event store is hard to reconcile with a right-to-be-forgotten request. The event-sourcing pattern's documented mitigation is to keep personally identifiable data out of the event stream itself, referenced by an opaque identifier, or to use crypto-shredding, encrypt each subject's data with its own key and delete the key to make the data unrecoverable without rewriting history.
Three tiers, three jobs, hot for what the agent needs right now, warm for what it should be able to find, cold for what it might need to prove later.
Step 3: Retries, Circuit Breakers, and Budget Guards
This is the section where a harness earns its keep, because it's also the section where the public evidence is thinnest. The clearest signal of unmet demand anywhere in this topic's search results is a Reddit thread about an agent that burned through a meaningful retry bill before anyone noticed, outranking a stack of short vendor blog posts on the same topic. None of those posts, vendor or community, ship a tested implementation. This section does.

Three failure classes need three different tools
Retry expects an operation to eventually succeed. A circuit breaker exists because retrying doesn't make sense for calls that are likely to fail, it prevents an application from hammering a resource that's already down. A budget watchdog catches a third thing neither of the first two handles: a loop that isn't failing at all, just running longer than it should. Microsoft's Retry pattern documentation is explicit that combining retry and breaker logic is normal, but warns against layering retries at multiple levels of a call stack, since each layer multiplying the others can turn one failure into dozens of calls.
Retry policy
Rate-limit handling differs by provider in ways worth knowing before you write a generic retry wrapper. OpenAI's rate-limits documentation instructs treating a Retry-After header as a minimum wait, adding a small random delay on top so multiple clients don't retry in lockstep, and disabling nested SDK-level retries if you're already retrying yourself, layered retries again.
Anthropic's rate-limits documentation confirms its errors also carry a retry-after header in the normal case, but flags one sharp edge: hitting a monthly spend cap returns a 429 with no retry-after header at all, and retrying, including the SDK's automatic retries, will keep failing until the cap resets (as documented as of September 2026). That's a case worth a distinct code path, not a longer backoff.
Retry with jitter
Python — retry policy (retry.py):
import logging
from tenacity import (
retry, retry_if_exception_type, stop_after_attempt,
wait_exponential_jitter, before_sleep_log,
)
logger = logging.getLogger("harness.retry")
class RateLimitError(Exception):
"""Raised on HTTP 429 from the model or tool provider."""
@retry(
retry=retry_if_exception_type(RateLimitError),
wait=wait_exponential_jitter(initial=1, max=60), # honors Retry-After as a floor
stop=stop_after_attempt(6),
before_sleep=before_sleep_log(logger, logging.WARNING),
reraise=True,
)
def call_model(payload: dict) -> dict:
response = provider_client.chat(payload)
if response.status_code == 429:
raise RateLimitError(response.headers.get("retry-after"))
return response.json()This mirrors the shape both OpenAI's own cookbook and Anthropic's documented conventions point to tenacity's exponential-jitter wait combined with a bounded attempt count is the pattern OpenAI's own rate-limit guide uses as its worked example.
Circuit breaker
Azure's circuit-breaker pattern documentation lays out the standard state machine: Closed (requests flow normally; a failure counter resets on a timer), Open (calls fail immediately without touching the provider, for a cooldown period), and Half-Open (a small number of test requests are let through; success closes the breaker, any failure re-opens it). The Half-Open state exists specifically so a recovering service doesn't get immediately flooded the moment it comes back.
A minimal breaker
Python — circuit breaker (breaker.py):
import time
from enum import Enum
class BreakerState(Enum):
CLOSED = "closed"
OPEN = "open"
HALF_OPEN = "half_open"
class CircuitBreaker:
def __init__(self, failure_threshold=5, recovery_timeout=30, half_open_max=1):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.half_open_max = half_open_max
self.state = BreakerState.CLOSED
self.failures = 0
self.opened_at = 0.0
def call(self, fn, *args, **kwargs):
if self.state is BreakerState.OPEN:
if time.time() - self.opened_at < self.recovery_timeout:
raise RuntimeError("breaker open: failing fast")
self.state = BreakerState.HALF_OPEN
try:
result = fn(*args, **kwargs)
except Exception:
self.failures += 1
if self.failures >= self.failure_threshold or self.state is BreakerState.HALF_OPEN:
self.state, self.opened_at = BreakerState.OPEN, time.time()
raise
self.state, self.failures = BreakerState.CLOSED, 0
return resultWhen the breaker is open, Azure's own guidance endorses graceful degradation over a hard failure, return a cached response, or route to a smaller fallback model for the easy cases while the primary provider recovers. Be precise about what that fallback actually is, too: this is model-tier routing (send easy queries to a small model, hard ones to a large one, a pattern Anthropic documents directly), not automatic provider failover. No captured documentation describes a provider transparently failing an in-flight request over to another provider on your behalf, so don't design around that assumption.
Budget guards and the loop watchdog
The retry and breaker layers both assume something is actually broken. Loops are sneakier, every individual call can look perfectly reasonable in isolation while the aggregate quietly drains a budget. The arithmetic is worth spelling out once: a ten-step run with three tool calls per step, each one a fresh inference carrying the full conversation context, is thirty-plus separate calls, not ten. One practitioner's account frames this as the actual discovery mechanism for the problem, there's no real "before" signal; you find out after the bill arrives.
Total inferences = Steps × Tool calls per step
Estimated cost ≈ Total inferences × Avg. tokens per call × Price per token
Multiple independent threads agree on this: telling the model "don't repeat yourself" in the system prompt does not reliably work, because from inside a single call, a repeated action looks locally justified every time. One practitioner's taxonomy of the failure modes is genuinely useful — identical repeated calls, A-B-A-B alternating loops, and no-op loops where the agent keeps calling despite an unchanging result — and the fix reported to work is a watchdog that lives outside the model: track a hash of each (tool, arguments) pair, and once it repeats past a threshold, block the call and return an explanation instead of executing it again. The same practitioner reports that surfacing the explanation back to the model ("you've called this three times — try a different approach") gets it to pivot more often than you'd expect.
Enforce token and cost budgets the same way, outside the model, checked against ControlFlags on every step, not inside a prompt instruction. Anthropic's guidance on agent loops explicitly recommends a maximum-iteration stopping condition for exactly this reason: control has to live somewhere the model can't reason its way around.
Idempotent side effects
A refund or ticket-creation call that fires twice because of a retry is the nightmare case for a support agent, and it's also the easiest one to prevent. Stripe's idempotency design is the template: attach an idempotency key to every side-effecting call, and store the result of the first attempt against that key regardless of whether it succeeded, a retried call with the same key gets the original result back instead of executing again. Store those keys in AgentState.idempotency_keys, keyed by the tool call they belong to, so a retry at any layer of the stack is safe by construction.
The failure-mode matrix
| Failure | Detection | Response policy | State action | Eval signal |
|---|---|---|---|---|
| 429 rate limit | HTTP 429 + retry-after header | Exponential backoff with jitter, honor header as a floor | Increment retry count on the call's ledger entry | Retry-rate per tool flags a noisy dependency |
| 5xx provider overload | HTTP 5xx / server error | Retry with backoff, escalate to breaker after threshold | Increment breaker failure counter | Frequency of breaker-open events per provider |
| Tool call timeout | Client-side timeout exceeded | Cancel, retry once with the same idempotency key, then breaker candidate | Log tool_timeout event to transcript | Per-tool p95 latency regression |
| Invalid JSON tool arguments | Schema validation fails on the call payload | Reject before execution, return a structured error to the model | Increment malformed-call counter; no side effect performed | Tool-call validity rate in trajectory evals |
| Repeated identical tool call | Hash of (tool, args) seen ≥K times this run | Block the call, return an explanation instead of executing | Set loop_flag = true; terminate or reroute | Identical-call rate per session |
| Circuit breaker open | breaker.state == OPEN | Fail fast; serve a cached response or route to a smaller model | Log breaker_open event with timestamp | Time-in-open-state per provider |
| Budget exceeded | tokens_spent > budget or step_count > max_steps | Stop the run, trigger the kill switch, return a partial result | Set control.budget_exceeded = true | Budget-exhaustion rate flags an inefficient prompt or tool |
| Escalation handoff failed | Human-handoff API errors or times out | Retry the handoff once idempotently, then fall back to an async ticket | Set escalation_state = pending_retry, reuse the idempotency key | Handoff-failure rate against SLA |
| Context overflow | HTTP 400, "prompt is too long" | Trigger hot-tier compaction, then retry | Set context_compacted = true, log token count at failure | Overflow rate bucketed by conversation length |
Retry, breaker, and watchdog are three separate tools solving three separate failure classes, using only one of them leaves the other two failure classes completely uncovered.
Step 4: Evals That Gate Deploys

Why "good" isn't a stable property
An agent that passes its tests today can fail the same tests next week without anyone changing a line of application code, the model gets updated, the prompt gets tweaked, a tool's API changes shape, or the retrieval corpus drifts. One practitioner's account of testing his own agent captured the trap precisely: he wrote his own test cases, and the agent got very good at passing exactly those tests, which tells you almost nothing about how it performs on cases it wasn't tuned against. That's Goodhart's Law showing up in agent testing: a measure that becomes a target stops being a good measure.
Three layers, one practitioner framework worth adopting
A useful mental model that surfaced repeatedly in practitioner discussion splits evaluation into three layers, and we adopt it here as the article's eval structure:
- Deterministic invariants, cheap, fast, binary checks: did the agent only call allowed tools, is every tool call's JSON valid, did the run stay inside its cost and latency ceilings? These run on every call and catch mechanical failures before they reach anything more expensive.
- Scenario tests, including refusal cases, did the agent do the right thing on a curated set of situations, including the ones where the right answer is refusing to act? A support agent that always tries to help is a support agent that eventually issues a refund it shouldn't.
- Weekly production sampling, a rotating sample of real traffic, reviewed against the same rubric, because synthetic test cases drift away from what customers actually ask over time.
Treat review burden, how much human reviewer time each eval cycle costs, as a first-class metric alongside pass rate. An eval suite that's accurate but takes eight hours of manual review every week won't survive contact with a real release cadence.
Golden trajectories and replay
Version the entire runtime profile that produced a result, model version, prompt version, tool versions, and the state of the retrieval index, not just the prompt. When a production run fails, replay it into the eval dataset as a new test case; the dataset should grow from real usage, not stay fixed at whatever you thought to write on day one. LangSmith's evaluation documentation distinguishes scoring a single turn from scoring an entire thread, full multi-turn coherence, not just one response in isolation, and that distinction matters here, because a support conversation's quality is a property of the whole exchange, not any single message.
Judging with your eyes open
LLM-as-judge is useful and imperfect in specific, documented ways. The research behind MT-Bench found that a strong judge model can reach over 80% agreement with human raters, on par with human-to-human agreement, but the same study catalogs position bias, verbosity bias, and self-enhancement bias as real failure modes worth designing around, not edge cases. Where a check can be made deterministic instead, did the agent call the correct tool, is the output valid JSON, make it deterministic; save the judge model for the genuinely subjective calls, like whether a response was appropriately empathetic.
Wiring evals into CI
Run the eval suite on every pull request and block merges on a regression, the same discipline a unit-test suite gets in any other codebase. Braintrust's evaluation guide describes exactly this loop: promote from a playground to a locked experiment, automate it in CI, then score live traffic and pull interesting production traces back into the dataset. DeepEval's documentation shows a pytest-style integration and dedicated regression view supporting the same pattern with open tooling if you'd rather not depend on a hosted platform.
The OpenAI Evals platform is being deprecated — it becomes read-only for existing users on October 31, 2026, and is scheduled to shut down entirely on November 30, 2026, as documented in its own guide. The vendor-neutral tools above (or a hosted alternative with a longer runway) are the safer foundation as of this writing.
Support-vertical rubric
For a support agent specifically, score four things: resolution correctness (did it actually solve the stated problem), escalation correctness (did it hand off when it should have, not just when it got stuck), deflection, the share of conversations resolved without a human, and satisfaction, gathered separately. Geckoboard's guidance on First Contact Resolution is worth internalizing here even though it's not an AI-specific metric: FCR should always be paired with a satisfaction measure, because a system can technically "resolve" a high share of contacts while leaving customers unhappy with how it did it.
Evals stop being nice-to-have the moment "good" stops being stable, which for an LLM-based agent is immediately.
Step 5: Observability: Instrument the Loop

Three timescales, one ID
Observability, monitoring, and evals answer three different questions on three different clocks: observability tells you what happened in this run, monitoring tells you whether something is broken right now, and evals tell you whether the system is better or worse than it was last week. A single trace_id running through all three is what lets you jump from "this metric spiked" to "here's the exact run that caused it" to "here's the eval case we should add because of it."
OpenTelemetry's GenAI conventions
OpenTelemetry's GenAI semantic conventions for spans, at "Development" stability as of September 2026, so expect naming to still evolve, give you a standard shape for this instead of inventing your own. Each operation gets a span named after its type and model (chat {model}, execute_tool {tool.name}, retrieval), carrying attributes like gen_ai.operation.name, gen_ai.request.model, gen_ai.usage.input_tokens/output_tokens, and gen_ai.conversation.id to stitch every span in a thread together.
On the metrics side, gen_ai.client.token.usage and gen_ai.client.operation.duration give you histograms you can alert on directly, and agent-specific instruments, gen_ai.invoke_agent.duration, gen_ai.invoke_agent.tool_calls, exist specifically for the loop-level view a single model-call span can't give you.
A traced model call
Python — OTel instrumentation (tracing.py):
from opentelemetry import trace
tracer = trace.get_tracer("agent.harness")
def traced_model_call(model_name: str, call_fn, payload: dict) -> dict:
with tracer.start_as_current_span(f"chat {model_name}") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.request.model", model_name)
span.set_attribute("gen_ai.conversation.id", payload.get("run_id", ""))
result = call_fn(payload)
span.set_attribute("gen_ai.usage.input_tokens", result.get("usage", {}).get("input", 0))
span.set_attribute("gen_ai.usage.output_tokens", result.get("usage", {}).get("output", 0))
span.set_attribute("gen_ai.response.finish_reasons", [result.get("finish_reason", "")])
return resultWrap tool calls and retrieval calls with the equivalent execute_tool {tool.name} and retrieval spans, and the full loop becomes traceable end to end under one conversation ID.
The PII decision the spec makes for you
The specification itself flags this: attributes that carry actual message content, gen_ai.input.messages, gen_ai.output.messages, and similar are marked Opt-In specifically because they're likely to contain sensitive information.
For a support agent handling account details and payment questions, that's not a hypothetical; treat transcript content as sensitive by default, redact or hash it before it reaches a trace, and keep the entry point for that decision in mind alongside the broader injection and disclosure risks the OWASP Top 10 for LLM Applications catalogs, including sensitive information disclosure as one of its named categories in the 2025 edition.
Cost receipts, not just a total-spend number
Log enough per-run detail to answer "why did this run cost what it cost" after the fact, not just "how much did it cost." One practitioner's list of the fields worth logging per run, workflow, model used per step, how much context was loaded, how many retries fired, whether a tool loop repeated, whether a fallback model kicked in, maps almost directly onto the OTel attributes above, which is a large part of why instrumenting the loop this way pays for itself quickly.
Choosing a backend
Langfuse's documentation describes it as OpenTelemetry-native and self-hostable, which matters if you want to avoid another vendor dependency sitting between you and your own trace data. LangSmith bundles tracing with its own evaluators and annotation queues, trading some of that independence for tighter integration with the eval workflow from Step 4. Either is a reasonable choice; what matters more than which one you pick is that the spans exist at all, a worked observability implementation with MLflow walks through wiring a specific backend to this same instrumentation if you want the next level of detail.
Instrument the loop once, with a standard convention, and you get debugging, cost accounting, and monitoring out of the same spans instead of building three separate systems.
Step 6: Ship to Production

Guardrails as a separate call, not a shared one
Anthropic's own guidance recommends running content and safety screening as a separate model call from the one generating the response, one instance handles the query, a second screens it, and the two together outperform a single call trying to do both jobs at once.
Pair that with a tool-risk tiering scheme, the pattern OpenAI's guide to building agents documents: classify tools by how much damage a wrong call could do, and gate the riskiest ones (issuing a refund, deleting a record) behind stricter checks or human confirmation than a read-only lookup needs. Designing that tool surface carefully matters as much as the guardrail logic itself, designing tools that agents can use safely is worth reading alongside this section.
Escalation as a designed state, not an error path
Human handoff shouldn't be what happens when the harness gives up, it should be a state the schema names explicitly, the escalation_state field defined back in Step 1. Make the handoff idempotent using the same key mechanism as refunds, so a retried handoff attempt doesn't create two tickets.
And hand the human a context bundle, not a blank slate: a summary plus a pointer back to the full transcript, so the customer doesn't have to repeat everything they already told the agent. Klarna's own announcement of its support assistant is worth citing for exactly one detail here, it explicitly kept the option to talk to a live human available throughout, rather than treating automation as a replacement with no way back.
Rolling out gradually
Shadow new versions against replayed cold-tier traffic before they see a live customer. Move to a small percentage of real traffic once the breaker and budget guards are armed, not before. Scale up against a defined SLO set: p95 latency, resolution rate, escalation rate, deflection rate, and review burden from Step 4.
Define deflection and satisfaction operationally in your own documentation rather than assuming a universal definition exists, no neutral, vendor-independent source defines them precisely, and pretending otherwise just hides the assumption instead of stating it.
Sizing the infrastructure before you need to
None of the numbers below are benchmark results, we haven't run this exact system at scale, and printing a made-up throughput figure would be worse than not printing one. What this table gives you instead is what actually drives the size of each piece, so you can plug in your own traffic and get a real estimate rather than guessing:
| Resource | What actually drives its size | Where it's covered |
|---|---|---|
| Redis (hot tier) | Concurrent active threads × per-thread payload size, bounded by the hot-tier token budget | Step 2, tier budget table |
| Postgres (warm tier + state) | Summaries and resolved entities written per conversation, times conversations per day, times retention | Step 1 (state), Step 2 (warm tier) |
| Object storage (cold tier) | Full transcript size × conversations per day × your compliance retention window | Step 2, cold tier |
| Model spend | The loop-cost arithmetic from Step 3 — steps × tool calls per step × tokens per call, at your provider's current price | Step 3, cost calculator |
| Eval / CI minutes | Test-suite run time × how often it gates a deploy, plus weekly production-sample review time | Step 4, review burden |
Compliance and the security boundary
Give the harness's tool credentials least-privilege access, the refund tool shouldn't have the same reach as a read-only order lookup, and treat the append-only transcript from Step 1 as your audit trail by default, since it already gives you a record of what happened without extra logging work.
If you operate where the EU AI Act applies, its logging obligations line up naturally with a harness that's already instrumented this thoroughly. None of this replaces a real security review, the enterprise threat model for agentic AI covers memory poisoning, blast radius, and the broader threat surface a production deployment needs to account for beyond what this article scopes.
The runbook, briefly
Tie your on-call response back to the failure-mode matrix from Step 3: a p95 latency spike usually traces to a provider slowdown or a breaker stuck half-open; a falling deflection rate usually traces to a retrieval or memory regression, not a model regression; a rising escalation-failure rate usually means the handoff idempotency key logic needs a look before anything else.
Build vs. Adopt: When Not to Build the Harness Yourself
Everything above is framework-agnostic on purpose, but that doesn't mean building it from scratch is always the right call. A framework like LangGraph gives you checkpointer and persistence plumbing for free, and that's a genuine time saving. The cost is real too: some abstraction opacity, and the risk of undocumented internal changes shipping underneath you, a concern raised independently by more than one practitioner working with framework checkpointers over time.

The decision rule that holds up: for a prototype, or a workflow with low stakes and no side effects, take the framework and its plumbing gladly. For a support agent issuing refunds under a strict budget, where you need audit-grade state and full control over retry semantics, own the loop yourself, or use a framework, but only one that exposes its internals cleanly enough that middleware can reach the seams this article customizes.
Anthropic's own agent-building guidance leans the same direction for teams starting out: begin with direct API calls, and only add framework or multi-agent complexity once it demonstrably earns its keep.
- tools have real side effects (payments, account changes)
- you're on the hook for an audit trail or regulatory review
- a runaway loop would be a real budget or safety incident
- you need quality that provably doesn't regress release to release
- it's a prototype or internal tool with no external users yet
- every tool is read-only — nothing to make idempotent
- a bad run costs a retry, not a refund or a compliance finding
- a framework's default checkpointing already meets your bar
If you're choosing between frameworks rather than choosing whether to use one at all, how the major agent frameworks actually compare is the natural next read, and if sessions, state, and memory as first-class framework primitives interest you specifically, Google ADK's approach to sessions, state, and memory is worth comparing against the schema this article built by hand.
The Production Readiness Checklist
Everything in this walkthrough converges into one artifact you can run before every launch and after every significant change: a pre-flight checklist. The thresholds below are the starting values from our design for the support-agent example, author choices, not industry constants, so tune them to your traffic and risk profile. Run the four phases in order, because each phase assumes the previous one passed.
Phase 1: State and memory foundations (Steps 1–2)
Production State Checklist
Ensure your state architecture meets these verification criteriaAgentState validates every checkpoint; new fields default to None so old transcripts deserialize via the schema_version upcast.Phase 2: Resilience armed (Step 3)
Resilience & Guardrails Checklist
Verification criteria for agent execution safety and recovery429 / 503 / timeouts retry — exponential backoff with jitter, a stop_after_attempt cap, and SDK-internal retries disabled.issue_refund with the same idempotency key cannot result in a double-refund.Phase 3: Evals gate deploys (Step 4)
Evaluation & Quality Checklist
Verification criteria for agent performance, safety, and operational costPhase 4: Production posture (Steps 5–6)
Observability & Operations Checklist
Verification criteria for tracing, security, alerting, and operational runbooksgen_ai.conversation.id.Twenty checkpoints is enough friction to catch the failure modes in the matrix above, and cheap enough to run before every deploy. If that feels heavy, remember what it replaces: the postmortem.
Common Mistakes
| Mistake | Why it fails |
|---|---|
| Treating checkpoints as memory | A checkpointer restores a thread's state after a crash; it isn't designed as a queryable long-term memory store, and using it as one leaves you with no summarization, no entity resolution, and no way to search across threads. |
| Vector-store-only memory | A raw similarity search over logs behaves like keyword search with extra latency and worse debuggability — it has no concept of which fact is current, contradicted, or stale. |
| Prompt-based loop prevention | Telling the model not to loop doesn't work, because each repeated call looks locally justified from inside a single inference; detection has to sit outside the model. |
| One global retry config for all tools | A refund API and a read-only lookup don't fail the same way or carry the same risk from a retry; a single retry policy for every tool either under-protects the risky calls or over-delays the safe ones. |
| Unbounded context "because the window is 1M tokens" | A bigger window doesn't cancel out context rot — performance still degrades as irrelevant history piles up, window size or not. |
| Evals only as an offline one-off | A test suite that never re-runs against new model versions or new production traffic stops measuring anything current within weeks. |
| A total-spend dashboard as the only cost control | Knowing the total doesn't tell you which workflow, model, or retry loop caused it — you need per-run attribution to actually fix anything. |
| Shipping without a kill switch or budget cap | The first runaway loop in production becomes a bill, not a bug report, if nothing outside the model can stop it. |
FAQs : Building Production Agent Harness
What is an agent harness, and how is it different from an agent framework?
An agent harness is the runtime scaffolding that surrounds a model during execution: it drives the loop of model and tool calls, persists state, enforces budgets and safety limits, and collects evidence about every run. A framework is a set of building blocks you use to write the agent's logic. A common way to phrase it: a framework is for building an agent, while a harness is for running one, and in production you usually need both layers to work together.
How do you stop an AI agent from looping forever?
Practitioners consistently report that prompt instructions alone, such as telling the model not to repeat itself, do not stop loops, because each repeated call looks locally justified to the model. The working fix lives outside the model: detect loop patterns like identical repeated calls, A-B alternation, and no-op calls, cap retries per tool, and run a watchdog that terminates the run when thresholds are exceeded. Fixing the underlying state bug, usually failing to persist prior tool results so the agent can see them, removes most loops at the source.
Do retries alone prevent runaway costs, or do you need a circuit breaker?
Retries and circuit breakers solve different problems, and production systems typically need both. Retries with exponential backoff and jitter handle transient errors such as a single 429 rate-limit response. A circuit breaker stops calls to a provider that keeps failing, so a sustained outage or rate-limit spiral cannot drain your budget while the agent keeps trying. Budget caps and a kill switch complete the set, because some cost blowups come from loops rather than provider failures.
What belongs in an agent checkpoint, and is it the same as memory?
A checkpoint is a snapshot of an agent run's state at a step boundary: the transcript so far, tool-call results, control flags such as step and token budgets, and escalation status. Framework documentation treats checkpointing as short-term, thread-scoped recovery, meaning resuming after a crash or pausing for human review, not as long-term memory. Long-term memory lives in a separate store that survives across threads, which is why most production designs keep both a checkpointer and a durable memory store.
What is hot, warm, and cold memory in AI agents?
It is a tiered memory design borrowed from operating systems. The hot tier, for example Redis, holds the current conversation's scratch data with short TTLs. The warm tier, for example Postgres with pgvector, holds structured summaries and entities for semantic recall. The cold tier, object storage or a warehouse, keeps full transcripts for audit, replay, and evaluation datasets. Research on long-context models shows quality degrades as prompts grow, so the harness curates what re-enters the context window instead of replaying everything, and promotion policies summarize and move data downward as conversations age.
How do you evaluate a customer-support AI agent?
A practical setup layers three kinds of checks: deterministic invariants such as valid tool calls, JSON correctness, and cost and latency ceilings; scenario tests with golden trajectories, including cases where the correct answer is to refuse or escalate; and periodic sampling of production traces. Metrics usually combine task success or resolution correctness, escalation correctness, deflection rate, and review burden, which is the amount of human effort required before an output can be trusted. LLM-as-judge scoring is common but has documented position and verbosity biases, so deterministic checks are preferred wherever they apply.
Do you need a framework like LangGraph to build a production agent harness?
No. The core loop is a plain Python state machine around a model API, and building it yourself gives you full visibility into state, budgets, and retries. Frameworks can save real work by providing checkpointers, persistence, and abstractions, and community experience supports using them for that plumbing. The trade-off is opacity and occasional abstraction churn, so a reasonable rule is to prototype with a framework if you prefer, while keeping the critical seams, namely the state schema, budgets, retry policies, and side-effect idempotency, under your own control either way.
📋 Article Timeline & History
Successfully updated on September 25, 2026 with the latest details.
This article was originally published on September 21, 2026.
Was this article helpful?










[…] For the code-level implementation of those controls — state.py, memory.py, resilience.py, and observability.py — see our walkthrough on building a production agent harness in Python. […]
[…] For the surrounding harness that decides when a long-context agent is even allowed to go to production, typed state, retries, budget guards, evals, and OpenTelemetry spans, see our walkthrough on building a production agent harness in Python. […]
[…] retries, budget guards, evals, and tracing, becomes the production backbone; see our walkthrough on building a production agent harness in Python for that […]
[…] append-only transcripts, idempotency keys, and OpenTelemetry spans, see our walkthrough on building a production agent harness in Python, whose structure maps cleanly onto Article 12 logging and Article 15 […]
[…] For a step-by-step Python implementation of that harness, typed AgentState, hot/warm/cold memory tiers, retries, circuit breakers, evals, and OpenTelemetry instrumentation, see our walkthrough on building a production agent harness in Python. […]