Promotion Ceremony
The Deploy Gate, or: How a Model Earns Its Tag
Every model ships through the same ceremony, and the ceremony has teeth. A gate that only ever passes candidates is decoration; this one earns its keep by saying no.
The Five Promotion Contracts
- A prompted baseline must lose first: Before training begins, the task eval runs against a much larger generalist model (qwen3.5:35b-a3b) with a thoroughly engineered prompt. If the fine-tune cannot beat a prompt against hardware we already own, it has no reason to exist.
- The bar is judged on the worst seed: Evaluation runs across two held-out random seeds. The gate judges the worse number. A lucky seed cannot ship a model to production.
- Candidates never touch the live tag: New artifacts deploy strictly as
<model>-candidate. The gate evaluates them in staging, and only an unambiguous pass promotes to:live. The previous release stays behind a-prevrollback tag. - The bar never relaxes to go green: Loosening a threshold is a policy decision requiring human operator sign-off, never an automated fix.
- The gate is a contract the framework reads: Thresholds are declared in
vastops.toml. The evaluation producesmetrics.json, andtuneharness gatejudges it and exits non-zero on any miss. A missing metric counts as a failure rather than silence.
Interactive Deploy Gate Simulator
Test how the deploy gate judges real model candidates from our production ledger. Select a candidate below to inspect its stage-by-stage evaluation:
Six Real Model Refusals and What They Taught Us
The six models refused by the gate are the strongest evidence that the system works:
| Refused Candidate | Reported Validation Score | Gate Refusal Cause | Permanent Recipe Correction |
|---|---|---|---|
ops-doctor v3 | 97.2% overall accuracy | Lost 45 points on hard-negative SSH failures | Checkpoint selection switched from eval-loss to label accuracy. Fixed in v4. |
sign-fielder v1 | F1 0.910 (beat baseline 2.5x) | Missed signature recall bar by 2.6% short | Nested braces in JSON collided with grammar. Switched to line format in v2. |
edge-adapter 3B | F1 0.924, sig recall 97.9% | Failed 4 of 4 exact multi-party canaries | Genuine 3B capacity wall. Candidate remains refused; 7B selected instead. |
rank-32 lora | Trained cleanly in PyTorch | Failed edge inference deployment | Vendor documentation stated rank 32, but real engine limit was 16. Local validation added. |
sign-pii encoder | F1 0.998 on held-out split | 0.286 F1 on format-disjoint OOD slice | Surface-form memorization caught by format-disjoint split. Cascade pattern deployed. |
trim-sft 1.7b | Vocab trimmed from 151k to 23k | Output validity collapsed to 0.065 | Input tokens in eval prompts became UNK tokens. Vocab trimming requires recovery pass. |
Why Canaries Beat Statistical Aggregates
Ten canary inputs with exact expected outputs run before the statistical evaluation. Rare catastrophic regressions do not significantly move an n=86 aggregate, but they trip an exact expectation immediately.
This was proven when an edge-serving candidate cleared the statistical bar outright (F1 0.924, signature recall 0.979) while failing the identical multi-party canary four times out of four. The aggregate had averaged the miss away; the canary refused the model.