Providers¶
A provider is a system under test: the thing domarinn sends inputs to and grades the output of. Every suite lists at least one under providers:, and the run matrix is providers × prompts × tests × repeats.
A provider is selected by its type. Five types exist:
type |
What it is | System under test? |
|---|---|---|
exec |
An external command speaking the exec JSON protocol (any language). | yes |
anthropic |
Native Anthropic Messages API client. | yes |
openai |
OpenAI-compatible chat-completions client (OpenAI + any compatible gateway). | yes |
http |
An arbitrary templated HTTP endpoint. | yes |
embeddings |
An OpenAI-compatible /embeddings client. |
no — it powers the similar assertion |
Every provider has an id (used in results and cache keys) and an optional label. The remaining fields depend on type.
Source of truth:
crates/domarinn-core/src/exec_provider.rs,anthropic.rs,openai.rs,http_provider.rs,embeddings.rs, the shared networking innet.rs, and theProviderKindschema inconfig.rs.
Environment-driven config¶
Any string in a provider's configuration — including elements of an exec provider's command argv and env map — may contain a ${env:VAR} placeholder, resolved once at load time — handy for a per-developer endpoint or a per-environment gateway that shouldn't be committed. The full rules (:-default, the $${...} escape, and exactly which parts of the suite this covers) are documented once, in domarinn.yaml → Environment interpolation.
exec¶
The flagship provider, and the escape hatch for testing anything you can run as a process: it shells out to a command that speaks the exec JSON protocol. If your program can read JSON from stdin and write JSON to stdout, it is a provider — no Rust, no SDK.
| Field | Type | Default | Meaning |
|---|---|---|---|
command |
[string] |
– | The command and its argv. Elements may contain ${env:VAR} placeholders. |
env |
{string: string} |
{} |
Extra environment variables for the child. Values may contain ${env:VAR} placeholders too. |
timeout_ms |
integer | 60000 |
Per-call timeout in milliseconds. |
cache_salt |
string | (none) | Provider-level version pin for the program — set it when a rebuild should discard cached answers. See below. Distinct from a test's own cache_salt, which keys that test's cases instead; see caching.md. |
Wire behavior¶
For each call the provider writes one provider request to the child's stdin and closes it, then reads one JSON response from stdout:
- Request (domarinn → child stdin):
{ "domarinn": {"protocol": 1, "kind": "provider"}, "prompt"?, "vars", "params", "test": {"id", "tags"} }.promptis null / omitted when the suite has no prompts (the "self-input" case) — the provider works fromvarsalone. A text prompt is sent as{ "text": "…" }; a chat prompt as{ "messages": [...] }. - Response (child stdout → domarinn):
outputis the only required field; see the protocol reference for the full set. A stringoutputbecomes text; any other JSON becomes a structured output.usagefills token counts,cost_usdfeeds thecostassertion, andmetadatais retained as the raw payload. - Worth reporting even though all of it is optional:
empty_reason(so a refusal is diagnosed instead of scoring zero against every assertion),error.class(so a rejected credential is distinguishable from a crash),error.details(structured diagnostics that survive to the stored case), andmodel(so an alias that silently repointed is visible).
The child always receives DOMARINN_PROTOCOL=1 in its environment, plus your env. The full wire contract, exit-code rules, and worked Bash/Python examples live in protocol.md.
Caching, and when you need cache_salt¶
exec providers are cached by default, under the same one rule as every other call: the key is a hash of the request — the command, its args, the protocol document on the child's stdin, and a digest of the declared env. It says nothing about the program's bytes, so an entry written on one machine is reusable on every other: a fresh clone, a different checkout path, a rebuilt binary and a different working directory all key identically.
The price is that domarinn cannot tell one build of your program from the next, so set cache_salt when a rebuild should discard the old answers — a commit SHA, a release tag, or "$digest: src/**/*.rs". Forget, and a hit whose stored program digest disagrees with what is on disk warns; nothing is invalidated, because whether a rebuild matters is the suite's call.
Anything else that steers the program belongs in argv or env rather than in a salt, where the key can see it. ${env:VAR} drives those from the ambient environment while keeping them keyed; a variable the child reads without the suite declaring it is invisible to the cache.
Full details — the two salt levels, $digest:, and what the child's environment does and does not key — are in caching.md.
Error and retry classification¶
An exec call is treated as retriable when the transport itself failed in a recoverable way — a spawn failure or a timeout — or when the child reports {"error": {"retriable": true}} in its response. A non-zero exit, unparseable stdout, or a child error with retriable: false is fatal. Retries follow the suite's runner.retries policy.
- id: assistant-fast
label: "assistant (fast model)"
type: exec
command: ["python3", "./assistant.py", "--model", "fast"]
# Environment for the child process only. Values interpolate from the
# ambient environment, so a per-developer endpoint stays out of git:
# ASSISTANT_ENDPOINT: "${env:ASSISTANT_ENDPOINT:-http://localhost:8080}"
env:
ASSISTANT_STYLE: "concise"
# A slow system under test needs a longer leash than the default.
timeout_ms: 30000
# Bump this when you edit assistant.py — the cache key hashes what the
# program is SENT, never its bytes, so an edited script otherwise keeps
# answering from stale cached results. See example 22 for the salt levels.
cache_salt: "dev"
anthropic¶
A native client for the Anthropic Messages API.
| Field | Type | Default | Meaning |
|---|---|---|---|
model |
string | – | The model id. |
base_url |
string | https://api.anthropic.com |
API base. |
api_key_env |
string | ANTHROPIC_API_KEY |
Env var holding the API key (sent as x-api-key). |
params |
object | {} |
Extra request-body params, passed through verbatim. |
request |
object | – | Transport overrides — auth scheme, headers, path, query, body overlay. See Customizing the request. |
cache_salt |
string | – | Cache pin. Change it to throw away every answer this provider has cached. |
Behavior:
- Calls
POST {base_url}/v1/messageswith headeranthropic-version: 2023-06-01. - Params pass through verbatim (
temperature,top_p,top_k,stop_sequences, …). Nothing is forced;model,messages, andsystemare set by the client and are authoritative over any same-named param. max_tokensdefaults to 4096 — the API requires it, so it is filled in only when yourparamsomit it.- System messages are extracted. A chat prompt's
system-role messages are pulled out and joined with blank lines into the top-levelsystemfield; the rest becomemessages. A plain text prompt becomes a singleusermessage. - Parses the response by concatenating
textcontent blocks. Records tokenusage(cache reads and both cache-write TTLs included),stop_reason, and themodelthe API reports having served.cost_usdis computed from the built-in rate table; seepricingto override it. - A missing prompt is a fatal error (this provider requires a prompt).
- id: claude
label: "claude-haiku-4-5"
type: anthropic
model: "claude-haiku-4-5"
# ANTHROPIC_BASE_URL is what the vendor's own tooling honours, so pointing
# this at a gateway needs no edit here.
base_url: "${env:ANTHROPIC_BASE_URL:-https://api.anthropic.com}"
api_key_env: ANTHROPIC_API_KEY
params:
max_tokens: 512
# USD per million tokens. Only the fields you state are overridden.
pricing:
input_per_mtok: 1.00
output_per_mtok: 5.00
cache_read_per_mtok: 0.10
cache_write_per_mtok: 1.25
openai¶
An OpenAI-compatible chat-completions client. Works against the OpenAI API and any compatible gateway (vLLM, LiteLLM, Together, Ollama, …) via base_url.
| Field | Type | Default | Meaning |
|---|---|---|---|
model |
string | – | The model id. |
base_url |
string | https://api.openai.com/v1 |
API base — point this at any compatible gateway. |
api_key_env |
string | OPENAI_API_KEY |
Env var holding the API key (sent as a bearer token). |
params |
object | {} |
Extra request-body params, passed through verbatim. |
Behavior:
- Calls
POST {base_url}/chat/completionswith bearer auth. - Params pass through verbatim. No temperature is forced; only
modelandmessagesare set by the client. - A text prompt becomes a single
usermessage; a chat prompt's messages pass through with their roles. - Parses
choices[0].message.contentas the output,finish_reasonas the stop reason, andusage.prompt_tokens/usage.completion_tokensas token usage.cost_usdis computed from domarinn's built-in rate table for known models; setpricing:to override the rates or price a model the table does not know. - A missing prompt is a fatal error.
- id: gpt
# The label interpolates the same variable as `model`: it exists to name
# the system under test wherever results are displayed, and a run against
# qwen3:4b must not report itself as gpt-4o-mini.
label: "${env:OPENAI_MODEL:-gpt-4o-mini}"
type: openai
model: "${env:OPENAI_MODEL:-gpt-4o-mini}"
# Defaults to the real API, and redirects with one variable — the same one
# the vendor's own SDK honours, so a gateway, a proxy or a local Ollama
# needs no edit to this file.
base_url: "${env:OPENAI_BASE_URL:-https://api.openai.com/v1}"
api_key_env: OPENAI_API_KEY
# Passed through to the API verbatim. temperature 0 is not a guarantee of
# determinism, but it removes the largest source of run-to-run noise.
params:
temperature: 0
max_tokens: 256
The same shape works unchanged against a self-hosted gateway — only base_url and api_key_env differ:
providers:
# A self-hosted OpenAI-compatible gateway (vLLM, LiteLLM, Ollama, ...).
- id: local-llama
type: openai
model: llama-3.1-8b-instruct
base_url: http://localhost:8000/v1
api_key_env: LOCAL_GATEWAY_KEY
params:
temperature: 0.2
Pricing¶
anthropic and openai providers cost each call from a built-in per-model rate table, so cost assertions and the run-level cost figure mean something without configuration.
A model the table does not know reports no cost at all rather than a guessed one — the cost assertion keeps honestly saying "not reported", and the run warns once naming the id. A made-up number that silently passes or fails a budget is worse than a loud no-op.
Ids rarely arrive in their plainest form, so three shapes resolve to the same row before that gives up. A dated snapshot has its date stripped (claude-opus-5-20260315, gpt-4o-2024-08-06); a Bedrock or Vertex decoration is peeled off (us.anthropic.claude-opus-5, claude-haiku-4-5@20251001); and anything left over falls back to the longest matching model stem, which is what prices suffixed aliases and point releases nobody enumerated:
That fallback is deliberately narrow. A stem only earns a fallback entry when no differently-priced sibling shares its prefix, which is why gpt-5, o3, gpt-4o and gpt-4o-mini have exact rows but no fallback: gpt-5-pro, o3-mini, and the gpt-4o-* audio-pipeline models (gpt-4o-mini-tts, gpt-4o-mini-transcribe) would inherit a rate that is off by several times. Models whose published price varies with context length — OpenAI's current flagships — are left out for the same reason, since a single rate cannot express two. Those ids stay unpriced and warn, and a pricing: block is how you price one anyway.
Override the rates, or price a model the table has never heard of, with pricing (USD per million tokens, merged field-wise over any built-in row):
pricing:
input_per_mtok: 1.00
output_per_mtok: 5.00
cache_read_per_mtok: 0.10
cache_write_per_mtok: 1.25
That works just as well for a model the built-in table has never heard of — a negotiated rate, a preview model, or a fine-tune behind a gateway — as it does for overriding a known one.
An exec provider that reports its own cost_usd always wins: it is the only party that knows whether it hit a proxy, a batch endpoint, or a different model entirely.
pricing is not part of any request, so setting it does not invalidate a single cache entry — cost is re-derived on every hit at the current rate. Cost is not request identity.
Graders are priced too, and reported separately¶
pricing also works on a grader.provider block and on the embeddings provider, so the models doing the scoring are priced by the same table and the same override. What they cost is reported as grader_cost_usd, next to cost_usd rather than added to it:
cost_usdis what the systems under test cost. That is the number acost:assertion budgets, and the one a model-selection decision turns on. A grader's price must not move a budget gate on the model being judged.grader_cost_usdis what measuring them cost. On a suite scored by a larger model than it tests, this is the bigger of the two; merging them would hide that rather than report it.
Per-assertion, the same figure appears as AssertResult.cost_usd. An exec grader reports nothing — the child spends against whatever endpoint it chose and the protocol gives it no way to say so, and a zero there would claim custom grading is free.
Credential preflight¶
Before the first call, domarinn checks that every credential the run will actually read resolves to a non-empty value, and fails with exit 2 naming the provider and the variable if not. "Actually read" is the operative part: a grader key is only required when a rubric assertion survived your filters.
It also rejects one known-wrong credential shape — an Anthropic OAuth access token (sk-ant-oat…), which the Messages API rejects as x-api-key. That is a hard failure only against api.anthropic.com; against any other base_url it is a warning, because a gateway may legitimately accept it.
The complaint is about how the credential is presented, not about the credential, so it does not fire for a provider that sets request: {auth: bearer} — that is the supported fix, not a workaround. A provider with auth: none reads no credential at all, so nothing is checked for it.
Without this, a wrong grader key errors every case in the suite and exits 3, which reads as an infrastructure fault after burning the run's entire provider spend.
Customizing the request¶
anthropic, openai, and embeddings know how to shape a request for their vendor. request: controls the envelope that carries it — everything about who you say you are, where you send it, and what rides alongside the body.
| Field | Type | Meaning |
|---|---|---|
auth |
api_key | bearer | none |
How the credential from api_key_env is presented. Defaults to the vendor's own scheme. none sends nothing and requires no credential. |
path |
string | Replaces the endpoint path appended to base_url (/v1/messages, /chat/completions, /embeddings). Must start with /. |
query |
object | Query parameters. Sorted by name before sending. |
headers |
object | Headers added to the request, overriding the vendor's own by name (case-insensitively). |
body |
object | Fields merged into the body last, after the provider built it. |
Every value is a minijinja template rendered against env. They are rendered once, when the provider is built — not per case. A header that must vary per case is what type: http is for.
Presenting an OAuth token¶
The motivating case. An Anthropic OAuth access token is rejected as x-api-key and accepted as a bearer token:
- id: claude-via-oauth
type: anthropic
model: "claude-sonnet-4-20250514"
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
# An OAuth access token rather than a Console API key. It is read at call
# time from the environment, exactly like any other credential — never
# written into a suite.
api_key_env: CLAUDE_OAUTH_TOKEN
request:
# `x-api-key` (the Anthropic default) would be rejected. Present the same
# credential as `Authorization: Bearer <token>` instead.
auth: bearer
headers:
# What the OAuth surface expects alongside the token.
anthropic-beta: "oauth-2025-04-20"
# TWO ENV SYNTAXES, AND THE CHOICE MATTERS.
#
# ${env:VAR} resolved at load time, IS in the cache key
# {{ env.VAR }} resolved at call time, is NOT in the cache key
#
# A selector belongs in the first: two tiers must not share one cache
# entry, or the second silently replays the first's answers. A
# credential belongs in the second: two teammates holding different
# tokens are asking the same question, and keying it would hand them
# private halves of a shared cache.
x-tier: "${env:MODEL_TIER:-standard}"
x-tenant: "{{ env.TENANT_TOKEN | default('anonymous') }}"
# Query parameters, for a gateway that routes on one. Sorted by name
# before they are sent, so two suites writing the same pairs in a
# different order still share a cache entry.
query:
api-version: "2024-10-01"
# Merged into the body LAST — after the provider has built it. That is the
# difference from `params:`, which merges FIRST and is then overwritten by
# `model`, `messages`, and `system`. Those three are exactly the fields a
# gateway sometimes needs changed, and `params:` cannot reach them.
body:
metadata:
user_id: "eval-harness"
domarinn performs no OAuth flow of its own: it does not fetch, refresh, or cache a token. Keeping one valid is the caller's job — a wrapper script, a CI step, whatever already mints it. auth: bearer only decides how the token you supply is presented.
body: reaches what params: cannot¶
params: merges first, and model, messages, and system are then written over it. Those three are exactly the fields a gateway most often needs changed — a routed model name, an injected system prompt — so params: structurally cannot reach them. request.body merges last, and can.
It is a deep merge: an object value merges key by key, anything else replaces. Overwriting messages is possible and is almost always a mistake.
DOMARINN_PROVIDER_HEADERS¶
Headers merged into every HTTP-speaking provider, as a JSON object, without editing a suite:
For an environment that must add a header to traffic it does not own the suites for. A suite's own request.headers wins by name — the variable supplies a default, and a suite that named the header meant it. Values are templates, rendered exactly like a suite's, so a credential written {{ env.X }} is redacted from the cache the same way.
Malformed JSON, or an object whose values are not strings, is a hard error rather than a silent skip: it was exported because a gateway requires it, and a dropped egress header fails at the far end with no local evidence.
These headers are part of the cache key. Exporting the variable in CI and not locally therefore splits the cache in two, which is the cost of treating them as request content rather than as transport.
http¶
A generic provider for black-box HTTP systems. The URL, headers, and body are templated (minijinja) against the test vars and the rendered prompt, and the response is projected to an output via an optional expression.
| Field | Type | Default | Meaning |
|---|---|---|---|
url |
string (templated) | – | Endpoint URL. |
method |
string | POST |
HTTP method. |
headers |
{string: string} (values templated) |
{} |
Request headers. In the cache key, as a digest rather than the values, so two providers differing only in X-Model do not share entries while an Authorization holding two teammates' tokens still does. |
body |
JSON (templated) | (none) | Request body, sent as JSON. |
output_expr |
string | (none) | minijinja expression selecting the output from the response. In the cache key: an entry stores the projected output, so changing the expression re-asks. |
Templating context for url / headers / body: every test var by name, plus two views of the rendered prompt:
prompt— the prompt as one string. A multi-turn prompt (amessages:prompt, or a case with history) is flattened torole: contentlines joined by newlines. Prose only: a turn's tool calls and itsthinkingblocks are deliberately absent, because inventing a textual rendering for a tool call is exactly what invites a tool-eager model to imitate that syntax as text instead of emitting a real call.messages— the same turns structurally, a list of{role, content}objects — plustool_calls/tool_call_idon a tool-using transcript, which is the view that carries them:{{ messages | tojson }}embeds the array JSON-encoded, and{% for m in messages %}iterates it — the way a body template forwards a real conversation to an OpenAI-shaped API. Atemplate:prompt appears as the single user turn it becomes on the wire.
A test var named messages takes precedence: the structural view is only added when no var of that name exists, so a suite that already forwarded a hand-rolled conversation under that name keeps rendering — and cache-keying — exactly as before. Existing templates that never reference messages render byte-identically either way.
The response is exposed to output_expr as:
response.status # integer HTTP status
response.text # raw body string
response.json # parsed body (if it parsed), else null
response.headers # response headers as an object
output_expr result handling: a string becomes a text output; any other value becomes a structured JSON output. Without output_expr, the raw response text is the output. This provider reports no token usage, cost, or stop reason.
Caching. The key is the request this provider would send: the rendered method, url and body, plus a digest of the rendered headers — and the configured output_expr, which never goes on the wire but decides what a stored answer means, so editing it busts that provider's entries. Test vars are in the key by way of the templates they render into.
One input can still change what happens without changing the key, and it is worth knowing:
- The environment, depending on which syntax you use.
${env:VAR}resolves at load time, so the substituted value is keyed — use it for anything that changes the answer.{{ env.VAR }}renders per request and is keyed as a literal${env:NAME}placeholder — use it for credentials, where keying the value would give every API key its own private cache. A provider whose url, headers or body reference{{ env.X }}warns at startup naming the variable, because domarinn cannot tell a model selector from a token. See caching.md.
- id: support-api
type: http
url: "${env:SUPPORT_API_URL:-https://api.example.com/v1/assist}"
method: POST
headers:
content-type: "application/json"
# Credentials come from the environment at call time, exactly as with the
# model providers. Never write a secret into a suite.
authorization: "Bearer {{ env.SUPPORT_API_TOKEN }}"
body:
message: "{{ user_message }}"
locale: "{{ locale }}"
# Reach into the parsed body for the field that is actually the answer.
# Note `response.json`, not `response` — the latter is the envelope.
output_expr: "response.json.result.reply"
embeddings¶
An OpenAI-compatible embeddings client. It is not a system under test — the runner filters it out of the graded matrix. Instead it powers the similar assertion: the first type: embeddings provider in the suite is handed to the grader.
| Field | Type | Default | Meaning |
|---|---|---|---|
model |
string | – | The embeddings model id. |
base_url |
string | https://api.openai.com/v1 |
API base. |
api_key_env |
string | OPENAI_API_KEY |
Env var holding the API key (sent as a bearer token). |
params |
object | {} |
Extra request-body params, passed through verbatim. |
pricing |
object | built-in rate | Rate override. Only input_per_mtok is read — see below. |
Behavior:
- Calls
POST {base_url}/embeddingswith bearer auth and body{ "model": …, "input": <text>, …params }. - Reads the vector from
data[0].embedding. Cosine similarity between the output and the reference then drivessimilar. - Listing an
embeddingsprovider as a direct system under test is unsupported (the provider factory rejects it); it exists only to servesimilar. - Each
similarassertion embeds two strings (the output and the reference), and both calls are priced and reported as that assertion's grading cost. Onlyinput_per_mtokapplies: an embedding call has no output tokens and the endpoint reports no cache counters, so the otherpricingfields would price components that do not exist.
- id: embedder
type: embeddings
model: "${env:OPENAI_EMBED_MODEL:-text-embedding-3-small}"
base_url: "${env:OPENAI_BASE_URL:-https://api.openai.com/v1}"
api_key_env: OPENAI_API_KEY
See similar for the assertion this provider exists to serve.
Retry, timeout, and error classification¶
The HTTP-backed providers (anthropic, openai, http, and the embeddings client) share one classification of failures, in net.rs. This keeps behavior consistent across every network provider.
| Condition | Classification | Retriable? |
|---|---|---|
HTTP 429 (Too Many Requests) |
retriable | yes — honors Retry-After |
HTTP 5xx (server error) |
retriable | yes — honors Retry-After |
Other 4xx (bad request, auth, not found) |
fatal | no |
| Timeout / connection / request build error | retriable | yes |
| Other transport errors | fatal | no |
| Missing API key env var | fatal | no |
Details:
Retry-Afteris honored. On a429/5xxcarrying aRetry-Afterheader (delta-seconds form), that delay is used before the next attempt; otherwise the runner's exponential backoff applies.- Retries are opt-in. The default is no retries (
runner.retries.max = 0). Configurerunner.retriesto enable backoff (default initial 500 ms, max 8000 ms). A retriable error that exhausts attempts becomes a caseerror(exit code3); a fatal error becomes a caseerrorimmediately. - Default request timeout for
anthropic,openai, andhttpis 120 s; the embeddings client uses 60 s. Theexecprovider uses its owntimeout_ms(default 60000 ms). - API keys never appear in cache keys. The key hashes the redacted request: credentials live in headers, which the keyed envelope excludes structurally, and an
execprovider's declaredenventers as a digest — precisely becauseenvis a credential channel, so the values never reach the stored entry. What the key does cover forexecis the command and its args plus the document sent to the child.
The two failure classes map to ProviderError::Retriable { retry_after } and ProviderError::Fatal internally, which is what the runner's retry loop keys off.
Falling back to another provider¶
Retries answer a provider that failed to reply. fallback: answers a provider that replied and said no — or one that is simply not there this morning. A refusal is not a verdict about the prompt, and a gateway being down is not a regression in the system under test, but without somewhere to go both become exactly that: a failed case and a red gate.
# The gateway that is having a bad day. When it returns nothing gradeable,
# the identical request goes to `backup` rather than the case reporting a
# failure the model never had a chance to cause.
- id: primary
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
fallback: [backup]
- id: backup
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
The list names other providers in the same suite, tried in order. The cell still belongs to primary: the request is handed on verbatim, and whoever answers, the case is recorded against the provider you configured.
When a link hands off¶
Two kinds of outcome, both keyed on open string types so a value this build has never heard of round-trips rather than being swallowed.
An empty answer whose reason is in the trigger set. runner.fallback_on_empty_reason defaults to ["refusal", "content_filter"] — the two that mean this provider will not answer this, as opposed to this provider answered badly, which is a result and belongs to the grader. Set it to [] to hand off only on hard failures, or widen it to any reason the classifier can produce.
A call that failed outright, in one of six classes:
| Class | |
|---|---|
provider_auth |
a rejected or missing credential |
provider_unavailable |
5xx, connection refused, DNS |
provider_timeout |
no reply inside the timeout |
provider_protocol |
a reply that could not be understood |
provider_rate_limit |
429 after retries are exhausted |
exec_failed |
an exec child that did not run |
This is an explicit list rather than "any infrastructure error", for two reasons. cache_miss and cache_unavailable classify as infrastructure and must never hand off — that would turn an offline run into a live call. And provider_request is deliberately absent: a 400 for a malformed body is the suite's bug, and the next provider will reject it identically, so spending a second call to learn that is noise rather than resilience.
Four things a reader can rely on¶
- Never under
--cache-only. Offline there is no live answer to go and get, so a handoff could only replace a usable — if refused — replay with a cache miss. This is a mode check, not a property of the trigger list, so no future trigger can weaken it. - Never for a cell carrying a
latencyassert. Such a cell is forced to cache-disabled rather than cache-only, so rule 1 does not cover it, andlatency_msis the answering link's time in flight. Sinceprovider_timeoutis a trigger, primary-times-out then fallback-answers-fast would make{latency, max: 2000}pass. "Never worse" needs a matching "never falsely better". - The primary is reported when nothing improved on it. If every link is also a trigger, the case settles on the primary's own outcome rather than the last link's — so a configured
fallback:can never make a case different from one with no fallback at all. Without this,provider_digestwould churn for zero gain and the run document would diverge from a no-fallback run, which breaks the server's re-upload idempotency. - Chains are not followed. A fallback's own
fallback:is ignored when it is reached as one, which makes a cycle unconstructible rather than something to detect mid-run.domarinn validatewarns when a target declares one, so the rule is discoverable while you are writing the suite. It also errors on an unknown id, a self-reference, and anyfallback:naming — or living on — anembeddingsprovider, which is a grader helper and never a system under test. Three checks coverfallback_only:: it is an error on anembeddingsprovider (which forms no cells anyway) and an error when every provider carries it (a matrix that can never grade anything), and a warning when no other provider'sfallback:names afallback_onlyprovider, since it can then never run as a cell or as a handoff.
Two of the guarantees are re-checked at run time rather than trusted from validate, because the library is callable without it: duplicate provider ids are refused outright — the last one silently winning would hand cells the wrong provider — and a chain link naming its own provider is dropped as the walk builds, so fallback: [self] is an error you are told about and a loop that cannot form.
A fallback that fails to build — no credential in the environment, say — is warned about and dropped, not fatal. Fallback providers are excluded from the credential preflight for the same reason: failing a whole run over a provider nothing selected would make fallback: a liability to configure rather than cheap insurance. fallback_only: providers are excluded on the same argument and more strongly — they are never a cell, so a run can only ever reach one through a chain that may not fire.
What the results say¶
cell.provider_id stays the configured provider, whichever link answered. That is what keeps case_key stable, so an --against baseline still joins the same row and a suite does not silently re-partition its history the first time a gateway hiccups.
What actually happened is recorded rather than hidden:
| Field | |
|---|---|
CaseResult.answered_by_provider_id |
who replied, set only when it was not the configured provider |
CaseResult.fallback_attempts |
each link tried and passed over, in order, with its empty_reason or error_class |
CaseResult.provider_digest |
the answering link's fingerprint, not the configured provider's |
RunSummary.fallback_cases |
how many cases needed one |
All four are omitted at their defaults, so a run that never fell back serializes byte-identically to one written before the feature existed. A server older than this feature accepts such a run and then drops the fields permanently, because it re-serializes its own typed struct on ingest.
Where they surface, so a handoff is not something you have to go looking for in JSON:
| Surface | |
|---|---|
domarinn case <id> |
· answered by <id> (fallback) on the header, plus a fallback: tried … line naming each link passed over and why |
the run Markdown (--summary-md, ci-summary) |
an Answered by fallback row reading N of G cases, where G is the graded total |
ci-summary step outputs |
a fallback-cases output, for a job that wants to branch on it |
--against diff |
a fallback cases: N → M headline when the count moved |
| the web UI | the case drawer and the matrix cell's popover |
An exec assert's cache key includes the answering provider's id
The assert context carries whoever actually replied, not the configured provider — the same truthfulness rule provider_digest follows. An exec assert sends that id to its child as provider.id, and the request it sends is the cache key, so the same output can grade under two different cache entries depending on who answered it.
That is deliberate. A child is entitled to judge differently based on which system produced the text, and a key that hid the difference would replay one provider's verdict for another's answer. The cost is that a suite whose primary hiccups occasionally will hold two entries per graded exec assert. llm-rubric is unaffected: its key is the judge's own call, and the provider id reaches only the legacy-adoption probe.
Because the digest moves, the server's compare view classifies such a pair as ProviderChanged — which is what it is, and it does so whether the case passed or failed. The CLI's --against is coarser: it reports a delta (newly failing, newly passing, unchanged) and does not carry that axis, so a case answered by a fallback on both sides reads as Unchanged even though two different models produced it. When a comparison matters, read fallback_cases on both runs before reading the delta.
A run where every graded case fell back exits 2. The suite ran, but not against the system it names, and a green gate would say otherwise. A partial fallback stays green — that is the feature working. --no-fallback is the recommended posture for a gate that would rather fail on its primary directly.
Filters, and which one applies¶
Three mechanisms narrow which providers a run touches, and they are three different questions. Confusing them is how a suite ends up either paying for a backup it never wanted in the matrix or losing the resilience it thought it configured.
1. A test's only_providers / skip_providers — cells and chain candidates. The only one that reaches both. A test that excludes a provider must not reach it through another provider's back door, so the exclusion is re-applied to every link before the walk considers it.
2. --provider — cells only. It chooses which cells run, and a cell that runs is entitled to its whole chain: a run that silently lost the resilience you configured would look exactly like fallback not working. So domarinn run --provider primary runs only primary's cells, and each of them may still reach backup. Naming a fallback_only provider is a usage error (exit 2) with its own message rather than a silent empty run — the id is in the suite, so "no such provider" would be a lie; it simply has no cells to select among. The whole list is vetted: an unknown, fallback_only, or embeddings id refuses the run even beside valid ids, because a half-right list that silently shrank the matrix would green-gate a provider the job never measured. DOMARINN_PROVIDER sets the same thing when the flag is absent.
3. fallback_only: — matrix membership. Not a filter at all: a property of the suite, set where the suite is reviewed rather than at the invocation, and therefore the same for every caller. It applies before --provider, because a flag can choose among cells but cannot conjure one for a provider that declared it has none. It is also deliberately unaffected by --no-fallback, which disables chain walking — a flag about resilience must not change which systems a suite is measuring.
fallback_only: a provider that is reachable but never a cell¶
Adding a backup to fallback: is cheap; adding the provider it names is not automatically cheap. A second entry under providers: is a second column of the matrix, so a suite that graded N cells now grades 2N and a nightly job's bill doubles to buy insurance it hopes never to use.
# The only system under test. Two tests here, one provider, two cells.
- id: primary
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
fallback: [reserve]
# In the suite, out of the matrix. It answers when `primary` will not, and
# adds nothing to what the run costs when `primary` is having a normal day.
- id: reserve
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
fallback_only: true
fallback_only: true keeps the provider in the suite — built, reachable, ready the moment primary will not answer — and out of the matrix, where it would cost something every run. The flag sits on the provider's outer struct, so it is common to every type: and, like fallback: itself, cannot move a cache key.
Example 45 is the runnable version, and contrasts it with example 44's --provider: selection at the invocation versus membership in the suite.
Cost attribution, and what it still cannot see¶
A case's cost_usd is correct: it is re-priced at the answering provider's rate.
Per-provider rollups used to bill the primary for the fallback's tokens, because the server promoted cell.provider_id into its cases table and nothing else recorded who answered. At the current migration they are answerer-correct: a case stores answered_by_provider_id, the matrix response carries run-level provider_costs keyed by the answering provider, and each MatrixCell counts how many of its cases were fallback_answered. So spend attributes to the model that actually produced the tokens.
Two things stay keyed on the configured provider, and both on purpose:
- Matrix columns. A column is the system under test, and re-pointing one because a gateway hiccupped would re-partition a suite's history exactly when you most want it stable — the same argument that keeps
case_keyfixed.fallback_answeredis how a column says "some of this was not me". A consequence to expect rather than debug:provider_costsmay name a provider with no column at all — afallback_onlyreserve has spend and no cells, which is precisely what it was configured to be. - Rows ingested before the migration, or by an older server, which never recorded an answerer and therefore attribute to the configured provider. A mixed history under-reports the fallback for the older span; read
fallback_caseson the runs involved before treating a trend as real.
Residual gaps, all in the same direction and none fixed by the above, since only the answering link is summarized: the primary's own spend on a call it handed off is not counted at all, a retry on the primary does not reach retried_cases, and a cache hit followed by a handoff counts as a miss.
Prose refusals: runner.refusal_patterns¶
A refusal is only classified as one when the vendor sets that finish reason. A model that answers "I can't help with that request." in ordinary prose is, as far as every mechanism above is concerned, a normal answer — so it is graded, and it is cached.
runner.refusal_patterns is the opt-in for that case: a list of regexes, empty by default, and an output matching any of them is treated as an effective refusal.
Three things to know before using it:
- A false positive silently swaps in a different model. A pattern loose enough to match a legitimate answer hands that case to the fallback and reports a
ProviderChanged, not an error. Anchor the pattern; test it against outputs you expect to keep. - A
not-containsassertion is the deterministic alternative. If what you want is this case must fail when the model refuses, assert it. That is a verdict about the response, computed the same way every run, with no dependence on which providers happen to be configured. - The pattern is re-applied on every read, never written onto a cache entry, so editing one reclassifies what is already stored rather than requiring a purge. An invalid regex is a
validateerror.
Patterns compile once per run, and matching a JSON output uses its rendered form — a refusal that arrived as {"error": "I can't help with that"} is still a refusal.
${env:VAR} works here, and is not a secret channel¶
Interpolation runs over the raw document before anything is parsed, so fallback: ["${env:FALLBACK_PROVIDER}"] works like any other config string, with the usual :-default rules. fallback: itself — and fallback_only: beside it — sits on the provider's outer struct and never reaches a fingerprint or a canonical request, so no value of either can move a single cache key. Turning on a chain, or excusing its target from the matrix, invalidates nothing in an existing suite.
Do not read that as "so environment values are safe to put anywhere"
Everywhere ${env:…} reaches something that is part of the request — a model, an endpoint path, an exec argv — it is resolved at load time and the substituted value is in the cache key. That is correct for anything that must separate two runs' entries. It is wrong for a credential, which must not, or a shared cache quietly becomes a private one per API key.
A credential belongs in a provider's own {{ env.X }} template, which is withheld from the key. It must never be routed through a case var: vars are resolved long before a provider renders anything, so the value reaches the request in the clear, is keyed, is stored on the entry, and is published in CaseResult.vars. The full split is in caching.md.
Example 44 — a second provider answers when the first refuses¶
# yaml-language-server: $schema=../../domarinn.schema.json
#
# A second provider answers when the first will not.
#
# A refusal, or a gateway that is simply down, is not a verdict about the
# prompt — but without somewhere to go, that is exactly what it becomes: a
# failed case, a red gate, and a morning spent looking at a suite that was
# never wrong.
#
# `fallback:` is resilience, not routing. The cell still belongs to `primary`,
# so `case_key` is unchanged and an `--against` baseline joins the same row.
# What changed is recorded rather than hidden: `answered_by_provider_id` names
# who actually replied, `provider_digest` is theirs, and the run summary counts
# how many cases needed it.
#
# Chains are NOT followed. If `backup` declared its own `fallback:`, it would
# be ignored when `backup` is reached as one — which makes a cycle
# unconstructible rather than something to detect mid-run. `domarinn validate`
# warns if you write one by accident.
#
# When it will not fire, and these are guarantees rather than accidents:
#
# - under `--cache-only`, where a handoff could only trade a usable replay
# for a cache miss;
# - on a cell carrying a `latency` assert, where a slow primary followed by a
# fast fallback would make the budget *pass* on a provider that never
# answered;
# - when no link improved on the primary — then the primary's own outcome is
# what the case reports, so a configured fallback can never make a case
# worse, or different, than no fallback at all.
#
# For a CI gate, consider `--no-fallback`: a gate usually wants to learn that
# its primary is broken rather than get a green from a different model. A run
# where *every* graded case fell back exits 2 for that reason.
#
# `--provider primary` below is not incidental. It chooses which cells run, and
# deliberately does NOT strip the chain — a run that silently lost the
# resilience you configured would look exactly like fallback not working.
#
# That is selection, made at the invocation. When the backup should never form
# cells at all — no flag to remember, and the suite itself states what a run
# costs — mark it `fallback_only: true` instead; example 45.
#
# Run: domarinn run examples/44-provider-fallback --provider primary
version: 1
project: examples
suite: provider-fallback
providers:
# The gateway that is having a bad day. When it returns nothing gradeable,
# the identical request goes to `backup` rather than the case reporting a
# failure the model never had a chance to cause.
- id: primary
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
fallback: [backup]
- id: backup
type: anthropic
model: claude-haiku-4-5
base_url: "${env:CLAUDE_GATEWAY_URL:-https://api.anthropic.com}"
runner:
# The default, spelled out. These two mean "this provider will not answer
# this"; every other reason an output can be empty is a result to be graded,
# not a reason to ask somebody else.
fallback_on_empty_reason: [refusal, content_filter]
prompts:
- id: only
template: "Answer in one short sentence: {{ question }}"
tests:
# The first case draws the refusal and is answered by `backup`.
- id: policy/return-window
vars:
question: "How long do I have to return an item?"
assert:
- type: contains
value: "30 days"
# The second is answered by `primary` normally — which is the point of having
# two here. A run where *every* graded case fell back exits 2, on the same
# argument the all-skipped guard makes: the suite ran, but not against the
# system it names. A partial fallback is the feature working, and stays green.
- id: policy/return-window-again
vars:
question: "What is the return window?"
assert:
- type: contains
value: "30 days"
See also¶
- protocol.md — the exec JSON protocol wire format for
execproviders, asserts, and generators. - caching.md — the one key rule, every cache knob in one table, and when
cache_saltis needed. - assertions.md — how provider outputs are graded, and the budget assertions that read
usage/cost_usd. - grading.md — using
anthropic/openaiproviders as the LLM-rubric grader. - domarinn.yaml — the full suite schema (
prompts,tests,defaults,runner,cache).