There used to be no official guidance on how to design the engineering layer wrapped around a model (commonly referred to as the harness—handing tools to the model, accumulating history, and executing the commands it emits; see below). OpenAI’s “A model guide for the GPT-6 family” and its accompanying mechanism documentation systematically describe for the first time how this engineering layer should be designed across the model family. Another change worth calling out separately is memory—the biggest headache when an agent runs continuously for hours (what it has done, what to do next). GPT-6 bakes this directly into the API, sparing developers from building their own engineering scaffolding. What developers really need to think about is how to collaborate with an agent that works autonomously over long stretches: which decisions it can make independently, and which require pausing to ask. Plus three mechanisms that don’t interrupt the task to handle course correction, waiting, and delegation.
Many people previously assumed that the most critical challenge in tasks lasting hours or even days was model memory: no matter how large the context window is, it cannot hold hours of content. People tried all kinds of extra methods—sliding windows, external vector search, having the model summarize its own history—to ensure the agent remembered what happened. From an engineering perspective, however, manually stitching context rarely works out well. The solution presented in OpenAI’s official GPT-6 guide is to let the vendor handle memory: encrypted compaction shrinks and persists histories that grow too long, while persistent reasoning preserves earlier chains of thought for subsequent turns. steering, async tool calling, and multi-agent handle the other end: correcting, waiting, and delegating without stopping the task. This redraws the boundaries of responsibility between model, harness, and human (async tool calling, multi-agent).
The main point of this article is that the challenge of long-horizon tasks is shifting from “how to make it remember” to “how humans and models coordinate.” The vendor has taken over memory, but preventing the agent from drifting off course during non-stop execution, knowing when to step back, and knowing when to make the call are what ultimately determine whether the task succeeds.
To give an example, an agent task designed to run continuously for several hours frequently encounters three situations in production: the context length falls short, truncating information mid-run; raw logs from intermediate tool calls flood the window, obscuring earlier hypotheses; or the agent halts at an uncertain branch to ask a question with no one around to answer, hitting a deadlock that strands the task. Another scenario involves synchronously invoking a slow tool—such as running tests or pulling a container image—where the entire harness freezes on that step, idling in place whenever a network timeout occurs. These situations repeatedly led everyone down the same path: doubling down on context management.
In such scenarios, an intuitive reaction is to invest heavily in context management, using methods outside the model to preserve and trim state. Similar experiments were conducted in the earlier article, published September 23, 2026 (the “September 23 article”), titled “Key Decisions in Harness Design: The Same Switch Boosted One Model and Broke Another”, which turned three modules of the agent’s outer harness into toggles and measured them individually across a code-repair benchmark and a terminal-task benchmark. Two findings are particularly relevant here. First, among the five context management strategies tested, a staged combination (pruning history with rules first, then having the model summarize older content if the window remained too long) incurred the lowest cost across all tested combinations (Section 47 of the September 23 article). Second, a dedicated readback mechanism (storing pruned long text externally and providing a tool for the model to retrieve it later when needed) delivered near-zero gains—in fact, its net impact on average accuracy was negative (Section 49 of the September 23 article).
Two data points illustrate how fragile this protection is: the difference in average success rate between having context management on versus off was 35.7 percentage points under a 32k small window, but dropped to just 2.7 percentage points under a 128k window (Section 47 of the September 23 article). Quadruple the window size, and the protection virtually vanishes. Engineering efforts invested in in-house compaction depreciate with every model generation; the pruning depth calibrated a year ago based on window budgets and invocation costs (Section 65 of the September 23 article) must already be reevaluated.
Things only began to change with the arrival of the GPT-6 family. OpenAI’s documentation notes that tasks running on the GPT-6 family can span hours or even days. While the compaction feature itself appeared earlier, what the vendor did this time was integrate it alongside persistent reasoning directly into the Responses API: the server automatically triggers compaction upon hitting a threshold, outputting an encrypted state blob that claims to encapsulate the state and reasoning required to resume the task. Why do they dare to package it this way now, freeing developers from worrying about how to write compression algorithms? The vendor’s explanation is that model capabilities have fundamentally shifted: GPT-6 generation models can pick up a task from an encrypted compaction block, resilient even when details are lost during compression; only when the model itself resists drifting off course can compaction be implemented so aggressively. A finding from the September 23 article is inverted here: in that article, harness strategies accommodated the model’s quirks (weaker models survived only when the outer harness spoon-fed them execution plans); now, the compaction infrastructure bets that the model is robust enough, staking everything on model capability. In exchange for convenience, developers must accept a vendor black box. When long-horizon tasks fail, they mostly fail in the coordination process: drifting with no one to steer, questions left unanswered, or slow operations bottlenecking the entire run. Once models possess the capability to run continuously over long horizons, what truly requires thought is how to collaborate with them without stopping the task.
While the previous three mechanisms address how to keep going without stopping, this section discusses the other side of the coin: what must stop. These two aspects are independent yet complementary: the former governs pace, while the latter governs direction. If the manifest is ambiguous, the model may repeatedly ask whether to proceed, fragmenting a long task into countless small stops, or take the initiative to make irreversible changes and ruin the task outright. Chapter 2 of OpenAI’s guide introduces authorization discipline for prompts and skills, which serves as the exact prerequisite for autonomous execution of long-running tasks in Chapter 3. Only when authorization is laid out clearly in black and white do those three mechanisms have a safe arena to operate.
When defining outputs, the official guide mentions that while the model has the discretion to structure a technical summary, it must consult a human before altering project scope. Similarly, Codex (OpenAI’s coding agent) features a clear tiering system: before a human steps away, specify which independent tasks should continue running and which decisions must pause to await a response. In terms of model behavior, GPT-6 Astra tends to ask frequent clarifying questions. The official model usage guide provides ready-made prompts to address this: advance and complete all authorized, verifiable steps first, deferring human approval to the final step before outward-facing actions are dispatched (official model usage guide).
At the API level, the vendor provides three coordination mechanisms that operate without interrupting tasks, each answering a single question: with this in place, what no longer needs to halt and wait for human intervention? The first is mid-turn steering. It allows host applications to inject corrective instructions via the Responses WebSocket API while model reasoning or tool execution is underway. OpenAI explains that such updates are queued and applied server-side without canceling in-flight tool calls or rolling back already-completed operations. The advantage of queuing is that if external monitoring detects a trajectory drifting off course, you can chime in with a single instruction to correct it rather than redoing the entire history. The documentation also notes a physical constraint: queued inputs are valid only on the current active connection; if the network drops, any unprocessed corrective instructions are lost.
The second is async tool calling, aimed at harness authors. If you
build a tool using the tool definition declared in an API request—such
as a function that runs tests or scrapes web pages, taking tens of
seconds—the traditional approach is synchronous: the model issues a
call, the entire response hangs there waiting for your function to
finish and return results, and only then does the model continue. The
asynchronous approach, however, adds async: true to the
tool definition. The model immediately moves on to other work after
issuing the call, while your program executes the function in the
background; once the function completes, the result is sent back for the
model to use. If you want the model to decide when to wait, you can
provide a custom wait tool (backed by a task handle registry where each
background task receives an identifier). While this is a recipe, the
critical change is that this capability shifts the failure mode of slow
tools freezing the entire loop from a harness code problem into
something eliminated directly by API syntax. (async
tool calling)
The third is multi-agent delegation. GPT-6.1 Sol also introduces a
beta multi-agent workflow in the Responses API, offering six managed
orchestration actions including spawn, send, wait, and interrupt. The
primary agent can delegate an independent subtask to a sub-agent for
parallel processing. The official documentation notes that the default
concurrency setting for max_concurrent_subagents is 3,
which is also the recommended value. There is no limit on delegation
tree depth or total agent count. In multi-agent mode, context compaction
for the root agent and each sub-agent is completely isolated; a
sub-agent’s trial-and-error trajectory is visible only within its own
lifecycle, preventing it from polluting the mainline context. (multi-agent
delegation)
A mature agent behavioral contract generally permits read-only and exploratory operations freely, while requiring strict review for writes or irreversible mutations. What the vendor has done here is essentially standardize the behavioral contracts developers previously had to implement inside their own harness.
Yet this introduces another problem. According to the documentation, GPT-6 Astra is notably more sensitive to the literal wording of rule files, including skills files and AGENTS.md (instructions placed in the repository for agents to read)—a point the official documentation highlights explicitly. It reads all rule files loaded into its context; if any of them is ambiguous or overly restrictive, the model may pause mid-run to request confirmation because of that rule, preventing the task from finishing. OpenAI’s model usage guide provides a solution: make this explicit in the instructions—current user instructions take precedence over skill rules, and add a self-check requirement directing the model to cite the exact file and rule that blocked it whenever it pauses mid-run.
Translating these principles into code yields five steps, each balancing vendor-managed defaults with areas developers can customize. The first step is selecting a model tier based on intelligence and cost. This step is a financial calculation. The official documentation frames model selection as a tradeoff between intelligence and cost. GPT-6 Astra ($10.00 / 1M input tokens, $50.00 output, $1.00 cached input) is the most intelligent model in the family, suited for the most complex reasoning and logical decisions. GPT-6.1 Sol ($2.00 input, $10.00 output, $0.10 cached input)—roughly one-fifth Astra’s price—is comparably intelligent and well-suited for complex coding, research, and computer use tasks where the model directly operates a desktop interface. GPT-6 Luna ($0.10 input, $0.50 output, $0.01 cached input) targets speed and low cost, handling tasks like invoice processing, classification, and fixed-format summarization that are high-volume and repetitive.
This tier determines how much thinking the model performs: * Low: Suitable for simple and routine tasks * Medium / High: Suitable for relatively difficult tasks requiring judgment * Extra high / Max: Experimental, enabled only when cost-effectiveness is justified The API allows changing tiers dynamically mid-session without invalidating the established prompt cache—the caching layer that discounts repeated prefixes. A common pattern is to start at a default tier, switch to High for critical code or challenging problems, and switch back once finished. Note that the API does not support two consecutive configuration updates, nor does it allow configuration updates concurrently with automatic compaction (official reasoning documentation). In addition, Fast mode is available exclusively through the API, offering faster and more consistent responses at a higher price. Ultrafast reduces latency by accelerating token generation speed (unrelated to reasoning) and is currently supported only on GPT-6 Astra.
Step 2: Use official compaction, do not delete history, and
arrange cached content in the right order Do not delete
historical messages; use the official
compaction system. You can set compact_threshold for
automatic compaction, or invoke /responses/compact manually
within the harness. The result of compaction is what is called a
compaction item: an encrypted blob completely opaque to humans. OpenAI
claims this encapsulates the state and reasoning required to resume the
task; sending it back in the next request lets the model pick up where
it left off. The rule of thumb: if you plan to use
previous_response_id (the parameter where the server
maintains conversational history for continuity) to chain calls, do not
delete historical messages yourself—let the server prune them. As for
caching, the vendor claims prompt caching can save up to 95% in costs.
In terms of arrangement, place system prompts, specifications, and tool
interfaces at the very beginning, and specific tasks at the very end.
Long-term budgets should also factor in cache writes and long-context
surcharges.
Step 3: Make long-running operations asynchronous, or use a
wait tool Mark operations taking more than a few seconds as
async: true. The harness maintains a
task_handle registry (one task_handle per
background task), allowing the model to receive a task identifier
immediately after issuing a call and proceed with other work; at the
same time, provide a custom wait tool so it pauses only
when it actually needs the result. Once background execution finishes,
backfill the result to the originating call_id to prevent
synchronous blocking (see the link above for
async tool calling usage).
Step 4: Author skills and AGENTS.md Provencher from OpenAI’s Developer Experience team observed that while the model’s comprehension of instructional nuances has improved significantly, overly rigid step-by-step procedures can sometimes hinder performance. You can follow the four recommendations in “Rethinking skills and prompts for GPT-6 Astra”: 1. Keep skill descriptions concise and clear, specifying only when to trigger them while loading detailed specifics on demand 2. In AGENTS.md, clearly specify when it applies to a given document or test, explicitly exempting safe testing on disposable local data from human review 3. Establish clear decision boundaries to avoid pointless questions 4. Define ‘done’—what constitutes task completion: modifying code, spinning up the environment, verifying results, fixing errors—and list items that strictly mandate human review At the same time, clarify priority: user instructions take precedence over skill rules, ensuring an obscure rule does not derail a long-running task.
Step 5: Enforce mandatory halts in production, correct drift with steering Establish mandatory human-approval blocking points in production (writing, merging, sending), while granting reasonable autonomy for local exploration, coding, and testing. If drift is detected, send queued steering instructions over WebSocket to correct course without interrupting the context.
The blueprint in this guide is a product manual tightly coupled to the OpenAI API. Copying it wholesale amounts to outsourcing architectural decisions to a vendor’s marketing page. Instability in product tiers represents the first major risk. In the GPT-5.6 era, OpenAI established three tiers—Sol, Terra, and Luna—calling it an enduring tiering strategy. Yet with the release of GPT-6, Terra vanished; the community speculated that Sol took its place, but OpenAI never confirmed this. Model tiers shift alongside business strategies, meaning they cannot be treated as system interfaces with long-term semantic stability.
From an engineering implementation standpoint, three details stand
out. First, steering is not persistent: it operates over WebSocket, and
once disconnected, it is gone, with no official cross-connection
recovery. Second, once multi-agent is enabled, the server disables the
/responses/compact endpoint, shutting down the avenue for
compacting and manually pruning the entire conversation history as a
whole. Third, certain API parameters are mutually exclusive: for
instance, reasoning.summary (a toggle for reasoning
summaries that outputs an overview of the model’s thinking process) and
max_tool_calls (the maximum number of tool calls permitted
in a single response) cannot be specified together. The API also rejects
back-to-back configuration_update calls (in-session tier
parameter updates), and configuration updates are temporally mutually
exclusive with automatic compaction.
More critical than tiering is validation: this mechanism lacks independent ablation studies. The September 23 article already observed that the exact same toggle can have diametrically opposite effects across different models. Removing dedicated file tools and leaving only the command line increased the success rate and cut costs by more than half on Nemotron-3 550B, but dropped Mistral’s success rate by 23.2 percentage points (Section 27 of the September 23 article). The planning toggle (tested in the September 23 article by strictly requiring in the system prompt to draft an execution plan first, re-injecting the updated plan into the context with each step forward) boosted the success rate of a 30B small model from 13.60% to 25.20%—an 11.6 percentage point gain that represented its lifeline—yet yielded zero score improvements on two frontier models, offering only cost-containment benefits (Sections 43 and 45 of the September 23 article). All of these were tested outside the OpenAI family. For the GPT-6 family’s new mechanism, no comparable controlled experiments exist yet.
This leaves three pivotal engineering questions for validation. First, can this non-interruptive task mechanism be implemented purely in client-side code for open-source models or multi-cloud architectures, without relying on OpenAI’s hosted WebSocket and encrypted compaction? Second, compared to the custom, transparent combination strategy validated in the September 23 article, how much evidence is preserved and how much information is distorted when opaque, encrypted compaction items run through multi-day real-world trajectories? The former offers convenience at the expense of auditability, while the latter offers the reverse; the jury is still out (comparative data in Section 47 of the September 23 article). Third, within the GPT-6 family itself, if the same harness is applied across Astra, Sol, and Luna, will it replicate the cross-model performance inversions documented in the September 23 article (Section 39 of the September 23 article)? There is also a longer-term question: once future models are heavily trained on real-world coordination patterns, how much of these outer coordination mechanisms will still need to be engineered by humans (the open question at the end of the September 23 article, Section 71 of the September 23 article)? This guide provides no answers; it is also where empirical data in this space is most urgently needed.