---
title: "Debug a failed run"
description: "Read a typed failure reason, act on its typed fix, check liveness before believing a state, and run the Doctor."
---
A terminal failure is never an opaque code and never a log wall. Every one carries four things a person can act on — what happened, why, a suggested fix in plain language, and often a typed `fix_action` a surface can render as a button.
## Start with the run
```bash
curl -fsS "$WS/worker-runs/$RUN_ID" -H "Authorization: Bearer $OBOL_SESSION_TOKEN"
```
The response is the whole run page in one object:
| Field | What it answers |
|---|---|
| `run` | the run itself: state, inputs, budget, deliverables, pinned revision |
| `liveness` | whether `run.state` can still be believed |
| `environment` | the leases, and `holds_compute` |
| `failure` | the typed failure, ready to render |
| `output_contract` | the contract evaluation; `null` means not reported, never passed |
| `steps` | the timeline |
| `most_expensive_step` | where the money went |
| `control_lease` | who, if anyone, is attending |
Read `liveness` **before** rendering `run.state`. A verdict of `unknown` is not "running" — it means nothing has been heard from the run recently enough to believe its last report. A page that prints `running` for a run nobody has heard from in an hour is telling a comfortable lie.
## The failure object
```json
{
"reason": "budget_exhausted",
"what": "The run stopped because it reached its spending limit before it finished the checkout walkthrough.",
"why": "This Worker's run budget is $2.00, and step 7 (\"Summarise the failing step\") took the run to $2.00 after 6 minutes.",
"rule_reference": { "kind": "budget_limit", "id": "wrv_4Hm9tZ", "label": "Run budget ($2.00)" },
"suggested_fix": "Raise this Worker's run budget to at least $3.00, or switch it to a cheaper model.",
"fix_action": { "action": "raise_budget", "scope": "worker", "suggested_minimum_micros_usd": 3000000 },
"most_expensive_step": { "step_attempt_id": "sat_19", "label": "Summarise the failing step", "usd_micros": 1420000, "share_of_total": 0.71 },
"set_by": "control",
"retryable": false
}
```
`set_by` is always `control`. Model prose never sets a failure reason.
The figures in that example are **this Worker's own configured budget**, not an Obol rate. Obol publishes no price; a budget is a ceiling the workspace set.
## The closed set of reasons
| Reason | Whose problem | Typical fix |
|---|---|---|
| `network_destination_denied` | your configuration | `widen_network_policy` |
| `connection_credential_expired` | your configuration | `reconnect_connection` |
| `budget_exhausted` | your configuration | `raise_budget` or `reduce_model_cost` |
| `deadline_exceeded` | your configuration | `extend_timeout` |
| `approval_expired` | nobody answered in time | `extend_timeout` on the approval window |
| `output_contract_not_satisfied` | the job or the contract | narrow the contract or the instructions |
| `environment_class_unavailable` | Obol's staging, or your plan | `request_capacity` |
| `capacity_exhausted` | Obol's | `retry` |
| `concurrency_limit_reached` | your plan | wait, or move up the ladder |
| `lease_lost` | Obol's | `retry` |
| `sandbox_provisioning_failed` | Obol's | `retry` |
| `ambiguous_effect_requires_review` | inherent | a human decides what happened |
| `policy_denied` | default-deny working | a reviewed policy change, not a retry |
| `revision_not_runnable` | the Worker as published | edit the Worker and publish a new revision; never retried |
| `canceled_by_user` | expected | — |
| `outcome_never_reported` | Obol's | check what happened, then decide |
| `internal_error` | Obol's | `retry` |
| `worker_reported_incomplete` | the job itself: the Worker said it could not finish | read its stated reason; never retried or repaired into a completion claim |
| `model_unavailable` | your model provider (5xx, transport error, timeout) | `retry` after backoff |
| `model_rate_limited` | your model provider (429) | `retry` after the provider's backoff, or raise the connection's quota |
The **Whose problem** column is also on the failure itself as `fault_domain`. Control derives it from the reason with one fixed table: `obol`, `customer_configuration`, `vendor`, `user_action` or `task_outcome`. A 401, or a non-policy 403, from your model provider is `connection_credential_expired`. A deterministic refusal to build the run's plan is `revision_not_runnable` and is not retried.
`outcome_never_reported` is the one reason Obol writes on its own initiative, and it is deliberately uncomfortable to read. It means the run passed its deadline while Obol still had it in flight and nothing ever reported an ending, so Obol closed it rather than show you a run that finished long ago as still running. It is not `deadline_exceeded`: that one is the runtime saying time ran out, this one is Obol saying it never learned. What the run actually did — including whether it completed external writes — is not recorded, so it carries no fix button and is not marked retryable. Check the systems the Worker writes to before you start it again.
Nothing outside this list can appear. Two consequences worth knowing: an unsupported *capability* currently reports as `environment_class_unavailable`, because there is no `capability_unsupported` variant; and an exhausted plan allowance never appears here at all, because that refusal happens before a run starts.
## Acting on the fix
`fix_action` is a closed union discriminated on `action`, so a client dispatches rather than parses:
```ts
import { matchFixAction } from "@obol/sdk";
const view = await obol.workers.getRun(runId);
if (view.failure?.fix_action) {
const label = matchFixAction(view.failure.fix_action, {
reconnect_connection: (f) => `Reconnect ${f.connector_slug ?? f.connection_id}`,
raise_budget: (f) => `Raise the ${f.scope} budget`,
widen_network_policy: (f) => `Allow ${f.destinations.join(", ")}`,
extend_timeout: (f) => `Extend the ${f.timeout}`,
reduce_model_cost: (f) => `Try ${f.suggested_model ?? "a cheaper model"}`,
request_capacity: (f) => `Request ${f.environment_class} capacity`,
grant_worker_policy: () => "Review the policy grant",
retry: (f) => (f.after_seconds ? `Retry in ${f.after_seconds}s` : "Retry"),
});
}
```
`matchFixAction` is exhaustive over the variants it knows, so a new one stops it compiling rather than silently falling through.
The contract has a ninth variant, `ask_copilot`. Since 2026-09-23 every failure offers at least that: a failure with no in-product repair still points somewhere. The dashboard renders it. The SDK's `matchFixAction` does not handle it yet, and `parseFixAction` returns `null` for it, so fall back to the failure's `what`, `why` and `suggested_fix`.
## Unresolved effects
A run whose write could not be confirmed now **pauses** instead of failing: it stays `paused` with `pause_reason: "ambiguous_effect_review"`. An approver records what actually happened:
```http
POST /api/v1/workspaces/{workspace_id}/worker-runs/{run_id}/effects/{logical_action_id}/resolution
Content-Type: application/json
{ "resolution": "applied", "note": "PR #42 exists" }
```
`resolution` is `applied` or `not_applied`. Control hands the decision to the run first, and records it only once the run accepts it. If the action is not awaiting review, the answer is `409 effect_not_awaiting_review`. If the run cannot be reached, it is `503`, and nothing is recorded.
`unresolved_effects` lists steps whose effect could not be established — typically an ambiguous browser or desktop interaction that may or may not have gone through. Do not retry those blind. `ambiguous_effect_requires_review` exists precisely so a person decides, and a Worker's default of zero ambiguous-write retries exists so nobody retries a write whose outcome is unknown.
## Check the setup, not just the run
```bash
curl -fsS "$WS/workers-doctor?worker_id=$WORKER_ID" \
-H "Authorization: Bearer $OBOL_SESSION_TOKEN"
```
The Doctor runs typed diagnostics — the Worker definition, connections, session profiles, domain access, model availability, budget headroom, concurrency, environment classes, triggers — each with a status, a message, and where possible the same typed `fix_action`. `overall` is `pass`, `warn` or `fail`, and `blocking_count` is how many things stop a run today.
A workspace-scoped report covers entitlement, availability, model connections and concurrency. Pass `worker_id` for the rest: a green workspace report must never be mistaken for a green Worker, which is why the scope is on the report.
Doctor reports are **not stored** in this build. What comes back is the only copy; `report_id` identifies the response you are holding and cannot be fetched again. Re-running produces a new report, which is also the honest behaviour — a credential can expire a minute after a check passes, which is why the run-start path re-checks rather than trusting a stored pass.
## When a route answers 503
Workers services are imported lazily, and a route whose service is absent answers `503 {"detail": "Workers service unavailable"}` rather than taking the control app down. A 503 from that path is a deployment fact, not a transient error to retry into — the route recovers when the service exists, not when you call it again.
Starting a run has its own 503, and it means something different. As of 2026-09-10 the orchestration hand-off cannot assemble a run input, so `POST /worker-runs` and `POST /worker-demo-runs` are refused there rather than by control. Where the orchestration service is not reachable at all, control does not translate the transport failure into a typed refusal and you get a `500` instead. Neither is a failed run: a run that never started has not failed, and no failure object is written for it. See [Overview](/workers/overview).