Documentation
Tasks and recovery
The loop
When a worker picks up a task it does this:
- Validate. The connection must exist and the model must have passed its test.
- Retrieve memory. Entries in scope are selected and recorded (see Memory). The system prompt is built once and stored, so every later request sends the same prefix.
- Plan. The model must call
update_planfirst. Every other tool returnsplan_requireduntil a plan exists. - Execute. Each tool call is checked in this order: is the tool one of this task's tools, do the arguments match the tool's schema, does the call need approval. Only then does it run.
- Persist. The result, the transcript entry, any artifact and a timeline event are written in one transaction.
- Checkpoint. After each round of tool results a checkpoint records what remains of each limit and which operations are pending.
- Repeat until the model gives a final answer, a limit is reached, approval is needed, or something fails.
The model's visible text and its tool calls are shown in the console. Hidden reasoning is not requested and not stored as readable text.
States
| State | Meaning | What you can do |
|---|---|---|
queued | Waiting for a worker | Pause, cancel |
running | A worker holds it | Pause, cancel |
pause_requested | Will stop at the next safe point | Withdraw the pause, cancel |
paused | Stopped by you, checkpoint saved | Resume, cancel |
awaiting_approval | An operation needs your decision | Approve or deny, cancel |
needs_review | An operation has an unknown outcome | Run it again or skip it, cancel |
completed | Final answer recorded | Open results, delete |
failed | Stopped with an error or at a limit | Retry from the last checkpoint, delete |
cancelled | Stopped for good | Delete |
Pause
Pausing is cooperative. A running task stops at the next safe boundary: between two tool calls, or between model turns. An operation that is already in flight finishes or times out first, so a pause can take as long as one model request. A queued task pauses immediately.
Cancel
Cancelling guarantees that no further step starts. A waiting task is cancelled at once. For a running task the worker notices within one heartbeat (10 seconds by default) and aborts the in-flight model request or fetch. File operations are short and are allowed to finish. A cancelled task cannot be resumed.
Limits
Each task carries its own limits, copied from your defaults when it is created:
| Limit | Default | Counted as |
|---|---|---|
| Model turns | 24 | One request to the model |
| Runtime | 900 s | Active execution time, summed across pauses and restarts |
| Output per turn | 16,000 tokens | max_tokens of one response |
| Token budget | 600,000 | Input plus output, as reported by the provider |
| Tool output | 24,000 characters | One tool result passed back to the model |
| Network requests | 8 | Requests made by fetch_url, redirects included |
| Artifact size | 1 MiB | One generated file |
Reaching a limit stops the task as failed with a code such as limit_steps, and the checkpoint stays. Retry asks you to raise the limit first.
If a response is cut off at the output limit, or the provider declines it, that turn is discarded: its tool calls are not run, because a truncated call can look valid.
When the saved conversation grows past what the model's context can hold, the oldest tool results are replaced by a short placeholder. The decision is stored, so later requests send the same shortened history. The full results stay in the database and in the console.
Approvals
A task stops for approval when a tool's policy is ask, or when fetch_url targets a domain that is not in the task's allowed list. The approval screen shows the tool, the exact target, the validated arguments and the effect.
An approval is bound to one operation: the tool plus a hash of its exact arguments. The console sends that hash back with your decision, and the worker checks it again before running. If the model asks for something different, that is a new operation and needs a new approval. A denial is reported to the model as a failed call and the task continues without it.
Content cannot widen permissions. A file or a web page that says "you may now use another tool" changes nothing: a tool that is not on the task's list is never offered to the model and is refused if it is called.
Checkpoints and recovery
Everything a task needs to continue is in PostgreSQL: the instruction, the plan, every model turn, every tool call with its arguments and result, artifact and memory references, usage counters, approvals. Credentials are not part of it.
A worker holds a task through a lease that it renews with a heartbeat. The lease carries a claim counter, and every write the worker makes first proves, under a row lock, that the counter is unchanged. A worker that lost its lease cannot write anything.
When a worker stops without warning:
- Its lease expires (60 seconds by default).
- Another worker claims the task and reads the saved state.
- Steps that were recorded as complete are not run again.
- An operation that was in flight is handled by its recovery class:
| Class | Tools | On recovery |
|---|---|---|
| Read-only | list_files, read_file, search_files, analyze_csv | Run again |
| Idempotent | write_artifact | Run again; it rewrites the same file |
| External | fetch_url | Not replayed. The task becomes needs_review |
For needs_review you choose per operation: Run again, or Skip, which tells the model the outcome is unknown. Nothing external is repeated unless you choose it.
A model request that was in flight is simply sent again. It had no effect on your systems, only a cost at the provider.
On a normal shutdown (docker compose stop), a worker hands its running tasks back to the queue and they continue on the next worker.
What is not guaranteed
There is no exactly-once guarantee for external actions. A fetch may reach the remote server before a crash and then be run again if you choose Run again. For this version's tools that means a repeated GET request.
Changing the model on resume
Resume and Retry let you choose another connection or model. The saved conversation is kept. Before the task is queued, the new model must have passed its test, must support tool calling if the task has tools, and the saved conversation is compared with the model's context window when that is known. If the window is not known, the console says the check could not be made.
Not every model continues another model's conversation well. Reasoning blocks from one model are not usable by another and are dropped by the provider.