A terminal failure is never an opaque code and never a log wall. Every one carries four things a person can act on — what happened, why, a suggested fix in plain language, and often a typed fix_action a surface can render as a button.

Start with the run

The response is the whole run page in one object:
Read liveness before rendering run.state. A verdict of unknown is not “running” — it means nothing has been heard from the run recently enough to believe its last report. A page that prints running for a run nobody has heard from in an hour is telling a comfortable lie.

The failure object

set_by is always control. Model prose never sets a failure reason. The figures in that example are this Worker’s own configured budget, not an Obol rate. Obol publishes no price; a budget is a ceiling the workspace set.

The closed set of reasons

The Whose problem column is also on the failure itself as fault_domain. Control derives it from the reason with one fixed table: obol, customer_configuration, vendor, user_action or task_outcome. A 401, or a non-policy 403, from your model provider is connection_credential_expired. A deterministic refusal to build the run’s plan is revision_not_runnable and is not retried. outcome_never_reported is the one reason Obol writes on its own initiative, and it is deliberately uncomfortable to read. It means the run passed its deadline while Obol still had it in flight and nothing ever reported an ending, so Obol closed it rather than show you a run that finished long ago as still running. It is not deadline_exceeded: that one is the runtime saying time ran out, this one is Obol saying it never learned. What the run actually did — including whether it completed external writes — is not recorded, so it carries no fix button and is not marked retryable. Check the systems the Worker writes to before you start it again. Nothing outside this list can appear. Two consequences worth knowing: an unsupported capability currently reports as environment_class_unavailable, because there is no capability_unsupported variant; and an exhausted plan allowance never appears here at all, because that refusal happens before a run starts.

Acting on the fix

fix_action is a closed union discriminated on action, so a client dispatches rather than parses:
matchFixAction is exhaustive over the variants it knows, so a new one stops it compiling rather than silently falling through. The contract has a ninth variant, ask_copilot. Since 2026-09-23 every failure offers at least that: a failure with no in-product repair still points somewhere. The dashboard renders it. The SDK’s matchFixAction does not handle it yet, and parseFixAction returns null for it, so fall back to the failure’s what, why and suggested_fix.

Unresolved effects

A run whose write could not be confirmed now pauses instead of failing: it stays paused with pause_reason: "ambiguous_effect_review". An approver records what actually happened:
resolution is applied or not_applied. Control hands the decision to the run first, and records it only once the run accepts it. If the action is not awaiting review, the answer is 409 effect_not_awaiting_review. If the run cannot be reached, it is 503, and nothing is recorded. unresolved_effects lists steps whose effect could not be established — typically an ambiguous browser or desktop interaction that may or may not have gone through. Do not retry those blind. ambiguous_effect_requires_review exists precisely so a person decides, and a Worker’s default of zero ambiguous-write retries exists so nobody retries a write whose outcome is unknown.

Check the setup, not just the run

The Doctor runs typed diagnostics — the Worker definition, connections, session profiles, domain access, model availability, budget headroom, concurrency, environment classes, triggers — each with a status, a message, and where possible the same typed fix_action. overall is pass, warn or fail, and blocking_count is how many things stop a run today. A workspace-scoped report covers entitlement, availability, model connections and concurrency. Pass worker_id for the rest: a green workspace report must never be mistaken for a green Worker, which is why the scope is on the report.
Doctor reports are not stored in this build. What comes back is the only copy; report_id identifies the response you are holding and cannot be fetched again. Re-running produces a new report, which is also the honest behaviour — a credential can expire a minute after a check passes, which is why the run-start path re-checks rather than trusting a stored pass.

When a route answers 503

Workers services are imported lazily, and a route whose service is absent answers 503 {"detail": "Workers service unavailable"} rather than taking the control app down. A 503 from that path is a deployment fact, not a transient error to retry into — the route recovers when the service exists, not when you call it again. Starting a run has its own 503, and it means something different. As of 2026-09-10 the orchestration hand-off cannot assemble a run input, so POST /worker-runs and POST /worker-demo-runs are refused there rather than by control. Where the orchestration service is not reachable at all, control does not translate the transport failure into a typed refusal and you get a 500 instead. Neither is a failed run: a run that never started has not failed, and no failure object is written for it. See Overview.