XT.PT The API → This story
Analysis The API

The 403 that can arrive after GPT-6 Astra's tool call has already run

GPT-6 Astra's migration notes, and the misalignment_policy_violation error OpenAI documents for stopped agents, read as a design brief for anyone running tools through it.

OPenAI - ChatGPT

OpenAI's API changelog entry for September 3 announced GPT-6 Astra, model ID gpt-6-astra, as "our most capable model, built for the hardest end-to-end work." XT.PT covered the model's slowed development in August. The API documentation that shipped with it is more useful to builders than the launch framing, because it spells out three things: which request parameters stop working, what the model costs, and what happens when OpenAI's own monitor decides an agent run needs a human.

What migration breaks

The changelog lists the key changes plainly. GPT-6 Astra "does not support the none reasoning effort level," "does not support custom temperature or top_p values or log probabilities (logprobs)," and "Tool calling requires the Responses API." The migration guide gets specific: remove temperature, top_p and top_logprobs; on Chat Completions also remove logprobs; on Responses remove message.output_text.logprobs from include. Callers on none or minimal effort are told to "start with low and compare results." The model page lists reasoning.effort values of low, medium, high, xhigh and max.

Chat Completions is still marked Supported on the model page, so a plain text integration keeps working. Anything agentic has to move, and that turns out to matter for the monitoring described below.

The numbers from the model page: a 1,050,000-token context window with a maximum of 922,000 input tokens and 128,000 output tokens, and an April 30, 2026 knowledge cutoff. Standard pricing is $10 per million input tokens, $1 cached, $12.50 for cache writes and $50 output. "Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request." Fast mode is not available with EU data residency.

A monitor that reads the agent, not the user

The fourth item in the changelog is the new one. OpenAI's guide defines misalignment monitoring as a check on "whether an agent is properly interpreting the user's instructions in consequential contexts, such as transferring sensitive data, accessing sensitive data, or making destructive changes." It "reviews model reasoning and actions asynchronously and can stop a conversation when it identifies a potential issue."

The guide is careful about what a flag means: it "does not establish that the user violated a policy or that the agent acted contrary to instructions," and "Monitoring can miss issues or flag legitimate activity." OpenAI still tells developers to keep their own safeguards, "including human approval for consequential actions."

Which requests it can stop

Coverage is not uniform. It depends on the API and on how the request preserves context:

Responses API with persisted reasoning,     monitored; can identify continuations
  WebSockets, or OpenAI compaction          and block further execution
Responses API using none of those           monitored; webhooks can receive alerts,
                                            no automatic stop
Chat Completions API                        not covered by this system

The guide adds that "Configuring an alert webhook does not enable automatic stopping." So the stop only reaches a conversation the system can recognize as continuing. A stateless Responses integration gets alerts. Chat Completions gets neither, but it also cannot do tool calling with this model, which is where the consequential actions live.

The error, and what it does not undo

When the monitor blocks a request before streaming begins, the documented error is:

HTTP 403
type: invalid_request_error
code: misalignment_policy_violation

OpenAI says to match on the code, not the message, and warns that streaming clients "must also handle errors while consuming the stream, even after receiving output." The prescribed handling is to stop dispatching actions for that conversation and not retry automatically, to preserve request and response IDs and tool records, and to put the error in front of whoever owns the task. There is no way back in: "The API does not provide a general way to resume a conversation stopped by misalignment monitoring."

The sentence that should shape harness design comes next:

Because monitoring is asynchronous, an action may already have completed before monitoring identifies a concern. A stopped request does not undo earlier actions.

OpenAI, Misalignment monitoring guide

The alert side has the same honesty. Projects can subscribe to a safety.alert.created webhook, whose payload carries only an alert ID, then fetch the alert with a key holding api.safety.alerts.read. A request_paused value of true means "registering a safety block succeeded; this does not confirm that execution stopped or that earlier actions were reversed." The reason field "can be null, including for Zero Data Retention (ZDR) requests," and "Alert delivery and retrieval do not provide a complete audit history."

What it asks of a harness

Read together, the docs describe a stop that is real but late, partial by design, and final. Nothing in them lets a builder treat it as a safety net. What they imply instead, and this is XT.PT's reading rather than OpenAI's checklist:

  • Your tool log is the record. The alert may carry a null reason and is not an audit trail, so the application's own record of which tool calls ran is the only way to know what a stopped run had already done.
  • Destructive tools need their own gate. An action can complete before the monitor objects, so irreversible operations still need approval or a dry-run step on the application side, which is what OpenAI itself recommends.
  • Plan for a dead conversation. A stopped conversation cannot be resumed, so long-running agents need checkpoints outside the conversation if the work is to be picked up by a human.
  • Know which coverage row you are in. The choice between persisted state and stateless calls now also decides whether OpenAI can halt your agent or only tell you about it afterward.

One more line from the Astra guide is worth a harness owner's attention. The model "can be more sensitive to instructions contained in skills and other files, such as AGENTS.md," and the guide says: "We strongly recommend auditing skills and other files accessible to your model for instructions that could influence its behavior." The flip side is ours to note, not OpenAI's: a model that follows file-borne instructions more closely is likely also easier for a poisoned file to steer.

Primary sources: OpenAI API changelog, GPT-6 Astra model page, Using GPT-6 Astra, Misalignment monitoring, read 2026-09-15.

Corrections and source documents: contact the desk
Read next →
Read next
Pricing · 4 min

GPT-6 Sol costs what GPT-5.6 Terra did, and exactly what Claude Sonnet 5.5 does

The API · 4 min

Claude Sonnet 5.5 makes thinking: disabled a 400, one of five breaking changes