Documentation
Provider setup
PoietesCMD talks to models through three adapters. You can have several connections, and you choose the connection and model for each task.
| Kind | Adapter | API used |
|---|---|---|
| OpenAI | openai SDK | Chat Completions |
| Anthropic | @anthropic-ai/sdk | Messages, streamed |
| Local / OpenAI-compatible | openai SDK with your base URL | Chat Completions |
Model IDs are yours to enter. Nothing is hard-coded: after a connection has a working credential, Load models from provider fills the list from the provider's own models endpoint, or you can type an ID.
Credentials
There are two ways to supply a key. Neither ever sends it to the browser.
Environment variable (preferred). Set the key for both the api and worker services, then create a connection that reads it:
OPENAI_API_KEY=...
ANTHROPIC_API_KEY=...
PCMD_LOCAL_API_KEY=...
PCMD_PROVIDER_KEY_ANYNAME=...
Only these names can be selected: the three fixed ones and anything starting with PCMD_PROVIDER_KEY_. The console shows whether a variable is set, never its value.
Stored in the database. Paste the key into the connection form. It is encrypted with AES-256-GCM using PCMD_ENCRYPTION_KEY and bound to that connection. The console shows the last four characters afterwards. This option is unavailable when PCMD_ENCRYPTION_KEY is not set. If you lose that key, stored provider keys cannot be decrypted and must be entered again.
To replace a key, edit the connection and enter a new one. To remove it, delete the connection.
Testing a connection
Test sends two real requests with the model you pick:
- A minimal text request. If it fails, the connection is marked Failed and shows the provider's error: wrong key, unknown model, no network, rate limit.
- A tool-calling probe: the model is asked to call one small tool. The result is recorded as Tools: yes or Tools: no.
A connection is shown as Connected only after the first request succeeded with its current credential and address. Changing the key or the base URL clears earlier results. If a running task later gets an authentication error, the connection goes back to Failed.
A model that has not passed the test cannot be used for a task. A model without tool calling can only run tasks with no tools selected.
OpenAI
Create a connection of kind OpenAI with a key and one or more model IDs available to your account.
Requests use the Chat Completions API with max_completion_tokens. The Responses API is not used, so reasoning items are not carried between turns on reasoning models.
Anthropic
Create a connection of kind Anthropic with a key and model IDs such as claude-opus-5-5 or claude-sonnet-5-5.
How requests are made:
- Requests are streamed and the complete message is collected, so long outputs do not hit HTTP timeouts.
- No
thinking, sampling ortool_choiceparameters are sent. Current models reject non-default values for them. Reasoning runs at the provider default and hidden reasoning is never requested or displayed. - Each assistant turn is stored exactly as returned, including its signed reasoning blocks, and replayed unchanged. The system prompt and the tool list are identical on every request of a task. This keeps reasoning valid on models that check it and keeps the prompt cache warm.
- Against the official API, automatic prompt caching and eager tool-input streaming are enabled. With a custom base URL (a proxy or gateway) both are left out, because a proxy may not know those fields.
Options on the connection:
- Effort. Sends
output_config.effort. Leave it on Provider default for models that do not support effort. - Refusal fallback (beta). Off by default. When on, requests use Anthropic's server-side fallback (
fallbacks: "default"), so a request declined by a safety classifier may be answered by another Claude model chosen by Anthropic. The task log records which model served the turn. It is off by default because it changes which model runs your task, and because it could not be verified without a key when this was built.
If a model declines a request, the task stops with The model declined this request and the category the provider reported. Tool calls from a declined turn are never run. You can resume the task with another model.
Local or OpenAI-compatible endpoint
Create a connection of kind Local / OpenAI-compatible with the base URL of a server that implements /v1/chat/completions, for example:
| Server | Base URL from inside Docker |
|---|---|
| Ollama | http://host.docker.internal:11434/v1 |
| LM Studio | http://host.docker.internal:1234/v1 |
| llama.cpp server, vLLM | http://host.docker.internal:8000/v1 |
Outside Docker, use http://localhost:<port>/v1. A key is optional.
Things to know:
- The base URL is administrator configuration. The server contacts it directly. It is not subject to the
fetch_urlrestrictions, and those restrictions are not relaxed for it: a task still cannot fetch a private address. - Many local models do not support tool calling. The test tells you. Without it, a task cannot use tools.
- Local servers often do not report token usage. The task then shows "usage not reported" and the token budget cannot be enforced; the turn and runtime limits still apply.
- Context windows are not reported by these servers. You can enter one per model on the connection so resumed tasks can be checked against it.
What is sent to a provider
For each model turn the worker sends: the system prompt (rules, limits, the memory entries supplied to the task), the task instruction, the tool definitions, and the conversation so far, which includes tool results such as file contents and fetched pages. With a cloud provider that content leaves your server. With a local endpoint it goes to that endpoint.
Provider keys, other tasks, files the task did not read and memory entries that were not supplied are not sent.
Verifying a real provider
The automated tests use doubles. To check a real provider from the command line:
OPENAI_API_KEY=... pnpm smoke:provider openai <model-id>
ANTHROPIC_API_KEY=... pnpm smoke:provider anthropic <model-id>
PCMD_LOCAL_BASE_URL=http://localhost:11434/v1 pnpm smoke:provider local <model-id>
It makes one text request and one tool-calling request through the same adapters the worker uses, asks for the provider's model list, and prints what came back: stop reason, reply text, tool call, serving model, token usage. The key is never printed. It ends with Result: PASS or Result: FAIL and exits non-zero on failure.
In the console the same two requests are behind Settings → Providers → Test.