Promotion Ceremony

The Deploy Gate, or: How a Model Earns Its Tag

Every model ships through the same ceremony, and the ceremony has teeth. A gate that only ever passes candidates is decoration; this one earns its keep by saying no.

The Five Promotion Contracts

  1. A prompted baseline must lose first: Before training begins, the task eval runs against a much larger generalist model (qwen3.5:35b-a3b) with a thoroughly engineered prompt. If the fine-tune cannot beat a prompt against hardware we already own, it has no reason to exist.
  2. The bar is judged on the worst seed: Evaluation runs across two held-out random seeds. The gate judges the worse number. A lucky seed cannot ship a model to production.
  3. Candidates never touch the live tag: New artifacts deploy strictly as <model>-candidate. The gate evaluates them in staging, and only an unambiguous pass promotes to :live. The previous release stays behind a -prev rollback tag.
  4. The bar never relaxes to go green: Loosening a threshold is a policy decision requiring human operator sign-off, never an automated fix.
  5. The gate is a contract the framework reads: Thresholds are declared in vastops.toml. The evaluation produces metrics.json, and tuneharness gate judges it and exits non-zero on any miss. A missing metric counts as a failure rather than silence.

Interactive Deploy Gate Simulator

Test how the deploy gate judges real model candidates from our production ledger. Select a candidate below to inspect its stage-by-stage evaluation:

Active Candidate:
Final Gate Verdict

Six Real Model Refusals and What They Taught Us

The six models refused by the gate are the strongest evidence that the system works:

Refused CandidateReported Validation ScoreGate Refusal CausePermanent Recipe Correction
ops-doctor v397.2% overall accuracyLost 45 points on hard-negative SSH failuresCheckpoint selection switched from eval-loss to label accuracy. Fixed in v4.
sign-fielder v1F1 0.910 (beat baseline 2.5x)Missed signature recall bar by 2.6% shortNested braces in JSON collided with grammar. Switched to line format in v2.
edge-adapter 3BF1 0.924, sig recall 97.9%Failed 4 of 4 exact multi-party canariesGenuine 3B capacity wall. Candidate remains refused; 7B selected instead.
rank-32 loraTrained cleanly in PyTorchFailed edge inference deploymentVendor documentation stated rank 32, but real engine limit was 16. Local validation added.
sign-pii encoderF1 0.998 on held-out split0.286 F1 on format-disjoint OOD sliceSurface-form memorization caught by format-disjoint split. Cascade pattern deployed.
trim-sft 1.7bVocab trimmed from 151k to 23kOutput validity collapsed to 0.065Input tokens in eval prompts became UNK tokens. Vocab trimming requires recovery pass.

Why Canaries Beat Statistical Aggregates

Ten canary inputs with exact expected outputs run before the statistical evaluation. Rare catastrophic regressions do not significantly move an n=86 aggregate, but they trip an exact expectation immediately.

This was proven when an edge-serving candidate cleared the statistical bar outright (F1 0.924, signature recall 0.979) while failing the identical multi-party canary four times out of four. The aggregate had averaged the miss away; the canary refused the model.