LLM-rubric grading¶
The llm-rubric assertion grades a provider's output against a natural-language rubric using an LLM as the grader — an LLM judge, in the usual phrasing. Unlike ad-hoc "ask the model and read the answer" grading, domarinn's grader never parses a verdict out of prose: it forces the model to return a structured verdict and treats anything less as an error.
Source of truth:
crates/domarinn-core/src/grader.rsand theGrader/AssertKind::LlmRubrictypes inconfig.rs. This assertion is introduced in assertions.md.
The verdict shape¶
Every grader — regardless of provider — must produce the same three fields:
| Field | Type | Meaning |
|---|---|---|
reasoning |
string | A brief justification. Listed first in the schema so the model reasons before it decides. |
pass |
boolean | Did the output satisfy the rubric? |
score |
number | A [0, 1] graded score (clamped on ingest). |
The JSON Schema the grader enforces is exactly:
{
"type": "object",
"additionalProperties": false,
"properties": {
"reasoning": { "type": "string" },
"pass": { "type": "boolean" },
"score": { "type": "number" }
},
"required": ["reasoning", "pass", "score"]
}
The built-in grading system prompt is:
You are a strict evaluator. Grade the ASSISTANT OUTPUT against the RUBRIC. Return a boolean
pass, ascorein [0,1], and briefreasoning. Judge only what the rubric asks; do not reward effort.
The user message is assembled as:
Grader resolution¶
An llm-rubric assertion needs a grader. It is resolved in this order:
- The assertion's own
grader:block (a per-assert override), if present. - Otherwise the suite-level
grader:block. - If neither exists, the assertion is an error (fail-closed):
llm-rubric assertion has no grader configured (set suite grader or per-assert grader).
An errored assertion promotes the case to error and drives exit code 3, not 1. It is never a silent pass. See assertions.md.
Grader configuration¶
The grader: block wraps a provider plus grading options:
| Field | Type | Default | Meaning |
|---|---|---|---|
provider |
provider spec | – | The grader model. Only anthropic and openai are supported for grading. |
template |
string | built-in | Optional file:// override of the grading-prompt template, relative to the suite directory. It renders into the prompt the grader reads, and the request is the key — so editing it re-grades. (An exec assertion's program is named by command and pinned by cache_salt, for the opposite reason: it receives the question rather than being part of it.) |
verdict_mode |
string | forced |
How the structured verdict is obtained: forced (default) or auto (rejected at load — not implemented). |
The provider is a standard ProviderKind — but only the anthropic and openai shapes are valid graders. Any other provider type errors with grader provider type … is not supported for llm-rubric.
verdict_modeandtemplateare part of the grader schema. The implemented grading path always uses the forced structured-verdict mechanism described below (a forced tool call on Anthropic, a strictjson_schemaresponse on OpenAI);forcedis the default and the mode you should rely on.
Suite-level grader with a per-assert override¶
A suite-level grader: block, straight from a shipped example — the grader is a
different model family than the system under test, max_tokens is raised well
above the default, and the credential is read only by the grader (see
Provider-specific mechanics):
grader:
provider:
type: anthropic
model: claude-haiku-4-5
base_url: "${env:ANTHROPIC_BASE_URL:-https://api.anthropic.com}"
# The grader reads ONLY what this names. It does NOT inherit the provider's
# credential resolution — which fails asymmetrically and confusingly:
# completions succeed while every grade dies on 401, so the run looks like
# an infra fault rather than a credential one.
api_key_env: ANTHROPIC_API_KEY
params:
# Raised well above the default on purpose. With a thinking model, a 1024
# default can truncate the structured verdict — which is a fail-closed
# error, not a silent pass. A generous ceiling costs nothing (you are
# billed for tokens generated) and removes a whole class of flake.
max_tokens: 4096
# `forced` is the default and the only value the loader accepts. `auto` is in
# the schema but was never implemented, so `domarinn validate` rejects it
# rather than quietly forcing anyway.
verdict_mode: forced
timeout_ms: 120000
Wired into a suite that uses it as the default grader, plus a per-assert override:
providers:
- id: sut
type: openai
model: gpt-4o-mini
prompts:
- id: qa
template: "{{ question }}"
tests:
- vars: { question: "Explain TLS to a five-year-old." }
assert:
# Uses the suite-level grader above.
- type: llm-rubric
value: "Correct, uses a simple analogy, no jargon."
# Overrides the grader just for this assertion.
- type: llm-rubric
value: "Strictly under three sentences."
threshold: 1.0
grader:
provider:
type: openai
model: gpt-4o
Provider-specific mechanics¶
Anthropic grader¶
The grader calls the Messages API (POST {base_url}/v1/messages, anthropic-version: 2023-06-01) and forces a submit_verdict tool call:
toolscontains a singlesubmit_verdicttool whoseinput_schemais the verdict schema above.tool_choiceis{ "type": "tool", "name": "submit_verdict" }, so the model must answer through the tool.- The verdict is read from the
tool_useblock'sinput. No prose is parsed. max_tokensdefaults to 4096 when yourparamsomit it.- The API key comes from
api_key_env(defaultANTHROPIC_API_KEY); the base URL defaults tohttps://api.anthropic.com.
Extended thinking is rejected. Forced tool use is incompatible with extended thinking, so a grader whose params include thinking or reasoning is rejected up front:
grader params must not enable extended thinking: forced tool use is rejected when thinking is on. Remove
thinking/reasoning.
Truncation is a loud error. If the response stops on stop_reason: max_tokens, the verdict is considered truncated and the assertion errors:
verdict truncated (stop_reason=max_tokens); raise grader max_tokens
Raise the grader's max_tokens (via params) and re-run.
OpenAI grader¶
The grader calls chat completions (POST {base_url}/chat/completions) with a strict structured response:
response_formatis{ "type": "json_schema", "json_schema": { "name": "verdict", "strict": true, "schema": <verdict schema> } }.- The verdict is parsed from
choices[0].message.content(guaranteed to match the schema by strict mode). - The API key comes from
api_key_env(defaultOPENAI_API_KEY); the base URL defaults tohttps://api.openai.com/v1, so any OpenAI-compatible gateway works as a grader viabase_url.
A suite-level grader in this shape, from a shipped example — any OpenAI-compatible endpoint can be the grader, a local Ollama included (see example 33):
grader:
provider:
type: openai
# Redirects with the same variable OpenAI's own SDK honours, and the same
# one example 26 uses for a provider — here it moves the GRADER instead.
model: "${env:OPENAI_MODEL:-gpt-4o-mini}"
base_url: "${env:OPENAI_BASE_URL:-https://api.openai.com/v1}"
api_key_env: OPENAI_API_KEY
params:
# See example 29: a thinking-capable grader can truncate a structured
# verdict at a small ceiling, which is a fail-closed error, not a
# silent pass. Generous costs nothing extra — you are billed for
# tokens actually generated.
max_tokens: 4096
# `forced` is the default and the only value the loader accepts. `auto` is in
# the schema but was never implemented, so `domarinn validate` rejects it
# rather than quietly forcing anyway.
verdict_mode: forced
timeout_ms: 120000
Truncation is a loud error. If choices[0].finish_reason is length:
verdict truncated (finish_reason=length); raise grader max_tokens
Parameters pass through verbatim¶
Whatever you put in the grader provider's params is merged into the request body verbatim — top_p, top_k, max_tokens, and so on. No temperature is forced by domarinn. The only rejected params are the thinking-enabling ones on Anthropic, as above.
Scoring an llm-rubric¶
Given the verdict { pass, score }, the assertion's pass/fail is:
- With a
thresholdon the assertion — pass whenscore >= threshold. - Without a
threshold— pass on the verdict's booleanpass.
The reported assertion score is always the verdict's score (clamped to [0, 1]); the verdict's reasoning becomes the assertion's reason. The case's weighted-mean score and pass/fail then follow the normal scoring rules.
# Binary: pass on the model's boolean judgment.
- type: llm-rubric
value: "Does the answer correctly identify the bug?"
# Graded: require a high score, not just a yes.
- type: llm-rubric
value: "Rate faithfulness to the source on the rubric below. 1.0 = every claim supported."
threshold: 0.8
Timeouts and defaults¶
| Setting | Value |
|---|---|
| Grader call timeout | 120 s, or grader.timeout_ms |
Default grader max_tokens |
4096 |
| Default verdict mode | forced |
grader.timeout_ms covers exec assertions as well as the HTTP graders — the ceiling belongs to grading, not to a transport.
Non-2xx responses, transport errors, missing tool_use/content, and truncated verdicts all surface as grader errors (fail-closed), recorded as grader error: ….
What grading costs¶
Grader calls are priced from the same built-in rate table as the systems under test, and a grader.provider accepts the same pricing: override. The same applies to the embeddings provider behind similar, which spends two calls per assertion (the output and the reference).
The figure is reported separately from the run's cost:
| Where | Field | Meaning |
|---|---|---|
| Run summary | cost_usd |
What the systems under test cost. |
| Run summary | grader_cost_usd |
What grading them cost. |
| Each assertion | cost_usd |
What that one verdict cost. |
They are not added together on purpose. cost_usd is what a cost: assertion budgets, and a grader's price must not move a budget gate on the model being judged. It also stays honest about the common case where the grader is the more expensive model: a merged number would bury that.
A grader call is cached like any other request, and its cost is recorded with it — so a fully-cached run still reports what its grading is worth rather than dropping to zero, re-priced at today's rate. Which calls were actually paid for this time is visible per assertion, via cached. --no-grader-cache re-asks the grader while still replaying provider responses; see caching.md.
An exec grader reports nothing: the child spends against whatever endpoint it chose, and the protocol gives it no way to say so. A zero would claim custom grading is free.
Best practices¶
- Use a different model family for the grader than the system under test. A model tends to prefer its own outputs (self-preference bias). If your SUT is GPT, judge with Claude, and vice-versa.
- Prefer binary or clearly-anchored rubrics. A crisp yes/no question ("Does the answer cite at least one source?") is far more reliable than a vague "rate the quality." When you do use a
score, anchor the endpoints in the rubric ("1.0 = every claim is supported; 0.0 = a fabricated claim appears"). - Write discriminative rubrics. A good rubric separates good outputs from bad ones — if every plausible answer would pass, the assertion isn't testing anything. State what a failure looks like.
- Keep
max_tokenscomfortable. Becausereasoningcomes first, the model spends tokens before emittingpass/score. If you see truncation errors, raisemax_tokensrather than shrinking the rubric. - Let deterministic checks gate the grader. Put cheap deterministic assertions alongside the rubric so obvious failures short-circuit the paid grader call.
Full example¶
version: 1
project: docs-demo
suite: rubric-grading
grader:
provider:
type: anthropic
model: claude-haiku-4-5
api_key_env: ANTHROPIC_API_KEY
params:
max_tokens: 2048 # pass-through; raises the truncation ceiling
verdict_mode: forced # default; shown for clarity
providers:
- id: sut
type: openai # different family than the grader
model: gpt-4o-mini
prompts:
- id: support
template: "Customer says: {{ message }}. Write a support reply."
tests:
- id: tone/apology
vars: { message: "My order arrived broken and I'm furious." }
threshold: 0.75
assert:
# Cheap gate: short-circuits the grader if it fails.
- type: not-icontains
value: "not my problem"
# Graded rubric with a scored threshold.
- type: llm-rubric
value: >
Rate the reply. 1.0 = acknowledges the problem, apologizes, and offers
a concrete next step (refund or replacement). 0.0 = dismissive or
blames the customer.
weight: 3
# Per-assert grader override + strict binary gate.
- type: llm-rubric
value: "The reply contains no promise the company cannot keep (no fake dates)."
grader:
provider:
type: openai
model: gpt-4o