Skip to content

Calibration

Calibration checks whether a step’s reported confidence actually means what it claims, empirically, using this exact system’s own history. For the plain-language explanation of why this exists, see How confidence and calibration work. This page is the technical/config reference: bucketing, labelling, binning, and the calibration: block’s knobs.

Every step execution’s effective_confidence has been persisted since Phase 0, and per-step/per-run human feedback (step_feedback/run_feedback) plus deterministic-check failures (Phase 1) have been accumulating. Phase 3 finally checks whether the number means what it claims: does a step that reports 0.75 confidence in a specific (step × agent × model × provider) configuration actually turn out correct roughly 75% of the time?

Bucketing. Marked step-executions are grouped by (step_name, agent, model, provider, prompt_hash, agent_version) — a library step (see Extending VectorStep) used across five pipelines feeds one bucket instead of five, and changing one step’s model resets only that step’s bucket. Fan-out branches (step_name like triage/0, triage/1) collapse into their parent step’s bucket rather than one bucket per branch index. prompt_hash/agent_version mean editing a step’s prompt template, or editing a Gateway agent’s agent.yaml/soul.md, also starts a fresh bucket — see How confidence and calibration work for the full explanation and what you’ll see in the UI when it happens.

Label precedence, per step-execution:

  1. Human — a resolved StepFeedback row for that step execution (correct → 1.0, partial → 0.5, incorrect → 0.0) — authoritative when present.
  2. Deterministic (D)pipeline_steps.deterministic_passed == False labels the step 0.0, for free, at scale. A passing check is not used as a positive label on its own — only failure is a strong-enough automated signal.
  3. Run-level fallback — the enclosing run’s RunFeedback.outcome, used only when neither of the above exists for that step execution.
  4. Otherwise the step-execution is excluded entirely — not counted as a 0, not counted toward N.

Binning, not curve-fitting. Rather than isotonic/Platt regression (which would pull in scipy/sklearn, a dependency this service otherwise has zero of), calibration uses simple fixed-width bins — default width 0.1 (10 bins across [0, 1]) — and reports each bin’s sample count and mean label. This is directly interpretable and matches the exact language calibration recommendations use: “runs scoring ~70% in this configuration are only 50% correct (40 runs).” A bin needs n_min (default 20) marked outcomes before it’s considered validated; nothing computed from an unvalidated bin is used to gate anything.

A step can opt its gate into using the bucket’s empirical accuracy instead of the raw self-report/verifier number:

- name: investigate
executor: gateway
executor_config: { agent: sre-investigation }
confidence_threshold: 0.75
calibration:
enforce: true
on_uncalibrated: proceed # or "escalate" — see below

When enforced and the step’s bucket/bin is validated, combined_trust is replaced with the bin’s mean_label before grounding’s min() and deterministic checks’ force-zero apply on top — the same confidence_threshold then decides on_low_confidence, no new threshold config. The TrustReport’s calibration block always shows the arithmetic: raw score, calibrated score, bin, n/n_min, so a calibrated escalation is never a mysterious abort.

When the bucket/bin has not yet accumulated n_min marked outcomes, on_uncalibrated decides the posture:

  • proceed (default) — combined_trust is left as the raw effective_confidence, unchanged; the run behaves exactly as it would with no calibration: block. The TrustReport still records “not yet validated, N=x/N_min” for transparency.
  • escalate — forces combined_trust = 0.0, driving the step’s existing on_low_confidence action. An explicit “no track record → a human checks” policy for high-blast-radius steps; not imposed as a universal default.

A step with no calibration: block is unaffected by any of this — same posture as Phase 1’s core invariant (see Grounding).

Blending in testing-stage marks (include_testing)

Section titled “Blending in testing-stage marks (include_testing)”

By default an enforced step’s bucket is built from stage: production history only — a brand-new step has no bucket at all until it’s promoted and accumulates real production runs, so enforce: true on a fresh step mostly just sits at not yet validated for a while. Set include_testing: true to let that step’s own gate draw on its stage: testing marks too, blended into the same bucket alongside any production ones:

- name: investigate
executor: gateway
executor_config: { agent: sre-investigation }
confidence_threshold: 0.75
calibration:
enforce: true
include_testing: true # default false

This is an explicit trade: the bucket validates faster, but it’s no longer a pure production track record — a change in behaviour you’re only exercising in testing could shift a bin’s mean_label before it’s ever seen real traffic. Reach for it on a step with little or no production volume yet where you still want the gate active, not as a default posture. TrustReport records include_testing alongside the rest of the calibration block, so a calibrated run is never ambiguous about which population produced its number, and the run detail page’s Trust panel calls it out explicitly whenever it’s set.

This only affects that one step’s own enforced gate. It has no effect on the Insights UI — /ui/insights/steps’s Production/Testing views never blend, regardless of any step’s include_testing setting (see Testing vs production stages) — so neither single-stage view there is the exact number an include_testing step’s gate is using; check that step’s own run detail page instead.

No persisted calibration curve. Calibration is still computed fresh from pipeline_steps/step_feedback/run_feedback on every request, the same way the Insights pages already recompute their rollups — there’s no fitted curve to migrate or invalidate. Prompt-versioning did add two columns to pipeline_stepsprompt_hash, agent_version — plus two small content-addressed registry tables that hold the recoverable text behind those hashes.

See samples/pipelines/trust-vector-remediation.yaml for a complete worked example combining critic/independent verifier modes, enforced grounding, deterministic checks, and calibration into a single trust-vector gate on a side-effecting step.

  • How confidence and calibration work — the plain-language walkthrough, including the worked five-step example and the full knob quick-reference table.
  • Promotion readiness — owner-defined criteria, including calibration-based readiness tiers, for promoting a pipeline out of stage: testing.
  • Grounding — the signal that caps combined_trust after calibration replaces the raw score.