Small-Model Training Substrate on Rented GPUs
Marketplace GPUs are cheap and hostile. We make them dependable.
An RTX 3090 rents for eleven cents an hour, pricing a 4B fine-tune at twenty-five cents. But the hosts misreport hardware, freeze mid-pull, and keep billing after training dies. tuneharness wraps marketplace compute with post-boot bf16 preflights, 30-minute idle watchdogs, append-only ledgers, and deterministic deploy gates that refuse substandard models.
Interactive CLI Walkthrough
A stdlib-only policy layer that verifies everything
Explore real terminal sessions recorded from actual marketplace runs. Click each workflow tab to inspect how the harness provisions, preflights, judges, and heals training jobs.
Defensive Engineering
Four invariants that turn strangers' rigs into reliable compute
Cloud orchestrators assume friendly datacenter hosts. Marketplace hosts are consumer rigs run by strangers. tuneharness assumes every host is lying until verified by renter-side code.
Post-Boot CUDA Preflight
An instance that passes nvidia-smi can still fail on real compute due to driver/toolkit mismatches. The harness runs an immediate torch.cuda.init() and bf16 matmul. Bad hosts are terminated and replaced in ten seconds, preventing the 105-minute silent idle burns typical of naive rentals.
Provider-Neutral Policy Seam
Exactly one module talks to vendor CLIs (vast.ai and RunPod). The policy layer (guardrails, quarantine, reputation scoring, cost ledger, watchdog, and deploy gates) never mentions the provider. Swapping marketplaces changes one adapter while all hard-won rules carry over untouched.
Money as a First-Class Failure Mode
The append-only ledger records create, stop, and destroy events with timestamps, rates, and reasons. Automated watchdogs sample GPU utilization on a cron and stop idle instances. Stop and destroy are asymmetric: automation may stop, but only the operator destroys.
Deploy Gate with Teeth
No model reaches production without beating its prompted generalist baseline on held-out data. Evaluations are judged on the worst of two seeds, tested on 10 exact failure canaries, and evaluated on format-disjoint out-of-distribution slices. Six candidate models have been refused to date.
Empirical Economics
Interactive GPU Marketplace Risk & Cost Calculator
Calculate the true cost of small-model training. Compare managed hyperscaler clouds against naive marketplace renting (with frozen pulls and silent idle burns) versus the guarded tuneharness pipeline.
Deterministic Deploy Gate
Where regressions die as candidates
A deploy gate that only ever passes models is theater. This gate earns its keep by saying no. Inspect real candidate promotion trials from our production history.
Empirical Grounding
The models, measured on held-out tasks
All evaluations are held-out and judged on the worst of two seeds. Baseline comparison is qwen3.5:35b-a3b (35B MoE generalist) running on our 24GB serving box.
| Model | Task | Size | Baseline Reference | Fine-Tune Result | Honest Slice / Caveat | Train Cost |
|---|---|---|---|---|---|---|
ops-doctor v4 | GPU-fleet action policy | 8B | Incumbent v2: 96.7% | 99.0% action acc, 1.2% missed | Same-simulator eval | ~$0.13 |
log-doctor v1 | Raw train.log triage (8 states) | 4B | Prompted 35B: 93% acc | 100% acc, 0% missed, 0.56s | Read 100% as eval exhausted | ~$0.35 |
ambie-eou v2 | Voice turn detection | 0.6B | Prompted 35B: 85% @ 524ms | 98.3% acc, 0.0% early-commit, 140ms | External OOD: 90.0% acc | ~$0.30 |
sign-fielder v3 | E-sign field placement | 4B | Prompted 35B: F1 0.421 | F1 0.922, sig recall 99.2% | Advisory OOD structural: 76.1% | ~$3.60 |
tool-caller sft | Schema-constrained tool calling | 1.7B | Prompted 4B: val 0.854 | val 0.874, valood 0.849 | 1.0000 valid with zero grammar tax | ~$0.26 |
Continuous Hardening
The Four Operational Feedback Loops
The interesting property is not any single feature. It is the loop: every operational failure becomes code the same day, and the pipeline gets harder to break every time it breaks.
Incidents to Code
An operational lesson may not live in human memory or documentation alone; it must land as a commit making the defect structurally impossible. Includes WAN flap grace windows, SSH proxy fallbacks, and autonomous OOM healing.
Protocolized Research
Training days start with automated knowledge checks for new architectures, toolchain updates, and live market rates. Research triggers action only if a candidate plausibly beats the incumbent by 2 points or halves cost.
Pipeline Models Babysit
The always-on monitoring ladder runs deterministic detection first, small specialist triage second (ops-doctor and log-doctor), and frontier LLM escalation only for ambiguous residue with capped spend budgets.
Tests Police the Tests
The harness maintains 100% branch coverage with mutation testing over money-critical paths. When a mutation gate broke silently, the fix made unverified execution a fatal exit, killing ten surviving mutants in cost accounting.