PoietesCMD

Documentation

Provider setup

PoietesCMD talks to models through three adapters. You can have several connections, and you choose the connection and model for each task.

KindAdapterAPI used
OpenAIopenai SDKChat Completions
Anthropic@anthropic-ai/sdkMessages, streamed
Local / OpenAI-compatibleopenai SDK with your base URLChat Completions

Model IDs are yours to enter. Nothing is hard-coded: after a connection has a working credential, Load models from provider fills the list from the provider's own models endpoint, or you can type an ID.

Credentials

There are two ways to supply a key. Neither ever sends it to the browser.

Environment variable (preferred). Set the key for both the api and worker services, then create a connection that reads it:

OPENAI_API_KEY=...
ANTHROPIC_API_KEY=...
PCMD_LOCAL_API_KEY=...
PCMD_PROVIDER_KEY_ANYNAME=...

Only these names can be selected: the three fixed ones and anything starting with PCMD_PROVIDER_KEY_. The console shows whether a variable is set, never its value.

Stored in the database. Paste the key into the connection form. It is encrypted with AES-256-GCM using PCMD_ENCRYPTION_KEY and bound to that connection. The console shows the last four characters afterwards. This option is unavailable when PCMD_ENCRYPTION_KEY is not set. If you lose that key, stored provider keys cannot be decrypted and must be entered again.

To replace a key, edit the connection and enter a new one. To remove it, delete the connection.

Testing a connection

Test sends two real requests with the model you pick:

  1. A minimal text request. If it fails, the connection is marked Failed and shows the provider's error: wrong key, unknown model, no network, rate limit.
  2. A tool-calling probe: the model is asked to call one small tool. The result is recorded as Tools: yes or Tools: no.

A connection is shown as Connected only after the first request succeeded with its current credential and address. Changing the key or the base URL clears earlier results. If a running task later gets an authentication error, the connection goes back to Failed.

A model that has not passed the test cannot be used for a task. A model without tool calling can only run tasks with no tools selected.

OpenAI

Create a connection of kind OpenAI with a key and one or more model IDs available to your account.

Requests use the Chat Completions API with max_completion_tokens. The Responses API is not used, so reasoning items are not carried between turns on reasoning models.

Anthropic

Create a connection of kind Anthropic with a key and model IDs such as claude-opus-5-5 or claude-sonnet-5-5.

How requests are made:

  • Requests are streamed and the complete message is collected, so long outputs do not hit HTTP timeouts.
  • No thinking, sampling or tool_choice parameters are sent. Current models reject non-default values for them. Reasoning runs at the provider default and hidden reasoning is never requested or displayed.
  • Each assistant turn is stored exactly as returned, including its signed reasoning blocks, and replayed unchanged. The system prompt and the tool list are identical on every request of a task. This keeps reasoning valid on models that check it and keeps the prompt cache warm.
  • Against the official API, automatic prompt caching and eager tool-input streaming are enabled. With a custom base URL (a proxy or gateway) both are left out, because a proxy may not know those fields.

Options on the connection:

  • Effort. Sends output_config.effort. Leave it on Provider default for models that do not support effort.
  • Refusal fallback (beta). Off by default. When on, requests use Anthropic's server-side fallback (fallbacks: "default"), so a request declined by a safety classifier may be answered by another Claude model chosen by Anthropic. The task log records which model served the turn. It is off by default because it changes which model runs your task, and because it could not be verified without a key when this was built.

If a model declines a request, the task stops with The model declined this request and the category the provider reported. Tool calls from a declined turn are never run. You can resume the task with another model.

Local or OpenAI-compatible endpoint

Create a connection of kind Local / OpenAI-compatible with the base URL of a server that implements /v1/chat/completions, for example:

ServerBase URL from inside Docker
Ollamahttp://host.docker.internal:11434/v1
LM Studiohttp://host.docker.internal:1234/v1
llama.cpp server, vLLMhttp://host.docker.internal:8000/v1

Outside Docker, use http://localhost:<port>/v1. A key is optional.

Things to know:

  • The base URL is administrator configuration. The server contacts it directly. It is not subject to the fetch_url restrictions, and those restrictions are not relaxed for it: a task still cannot fetch a private address.
  • Many local models do not support tool calling. The test tells you. Without it, a task cannot use tools.
  • Local servers often do not report token usage. The task then shows "usage not reported" and the token budget cannot be enforced; the turn and runtime limits still apply.
  • Context windows are not reported by these servers. You can enter one per model on the connection so resumed tasks can be checked against it.

What is sent to a provider

For each model turn the worker sends: the system prompt (rules, limits, the memory entries supplied to the task), the task instruction, the tool definitions, and the conversation so far, which includes tool results such as file contents and fetched pages. With a cloud provider that content leaves your server. With a local endpoint it goes to that endpoint.

Provider keys, other tasks, files the task did not read and memory entries that were not supplied are not sent.

Verifying a real provider

The automated tests use doubles. To check a real provider from the command line:

OPENAI_API_KEY=... pnpm smoke:provider openai <model-id>
ANTHROPIC_API_KEY=... pnpm smoke:provider anthropic <model-id>
PCMD_LOCAL_BASE_URL=http://localhost:11434/v1 pnpm smoke:provider local <model-id>

It makes one text request and one tool-calling request through the same adapters the worker uses, asks for the provider's model list, and prints what came back: stop reason, reply text, tool call, serving model, token usage. The key is never printed. It ends with Result: PASS or Result: FAIL and exits non-zero on failure.

In the console the same two requests are behind Settings → Providers → Test.

NextTasks and recovery