System Architecture
The Provider Seam and Renter-Side Verification
How tuneharness decouples marketplace dependencies from hard financial guardrails, validates hardware before training begins, and prevents money leaks.
+-------------------------------------------------------------------------+
| OPERATOR CONFIGURATION |
| - Max hourly rate ceiling ($0.35/hr) - Min reliability floor (0.98) |
| - Min download bandwidth (400 Mbps) - Auto-quarantine cooldown |
+------------------------------------+------------------------------------+
|
v
+------------------------------------+------------------------------------+
| TUNEHARNESS POLICY CORE |
| [Ledger] -> Append-only create, stop, destroy lifecycle tracking |
| [Leases] -> Label reconciliation (prevents duplicate instance leaks) |
| [Reputation] -> Time-decayed Beta posterior scoring per host |
| [Watchdog] -> 30-min idle sampler; stops unmonitored compute burn |
| [Gate] -> Deterministic deploy evaluation on held-out test splits |
+------------------------------------+------------------------------------+
|
+----------------------+----------------------+
| Provider Neutral Seam (Offer, Instance, State)|
+----------------------+----------------------+
|
+--------------------+--------------------+
| |
v v
+---------------------+ +---------------------+
| vast.ai Adapter | | RunPod Adapter |
| - CLI JSON wrapper | | - REST v1 Lifecycle |
| - Auto-retry parser | | - REST v2 Catalog |
+---------------------+ +---------------------+
1. The Seam: Where Vendor Stops and Policy Starts
A stdlib-only Python package wraps vendor CLIs rather than raw HTTP APIs, so vendor API drift is the vendor's problem. On top of that sit guardrailed search and provisioning, rsync-based code sync, tmux-managed training sessions, regex-parsed progress reporting, an idle watchdog, and a lifecycle ledger that records every create, stop, and destroy with price and reason.
The load-bearing decision in that stack is where the vendor stops and the policy starts. Exactly one module talks to the marketplace, and it does so through provider-neutral abstractions: offers, instances, prices, and lifecycle states. The guardrails, the ledger, the delivery verification, the watchdog, the self-healing loop, and the deploy gate never mention vast.ai or RunPod.
The practical consequence is that the provider is a dependency while the policy layer is the product. Swapping or adding a marketplace changes one adapter, and every hard-won rule about money, verification, and honesty carries over intact.
2. Trust Nothing: The Post-Boot Verification Pipeline
The core design principle is that marketplace hosts are untrusted until proven otherwise on every single boot:
- Fail-Fast Image Pull Watchdog:
BootProgressinspects host status messages during container download. If the pull status freezes for six minutes, or reports a known-fatal docker error, the attempt is killed immediately. A bad host costs 3-4 minutes instead of 25-30 minutes of poll-waiting. - Post-Boot CUDA Compute Preflight: An instance that passes
nvidia-smican still fail on real PyTorch execution due to stale host kernel drivers. The harness runs an immediate preflight:torch.cuda.init()and a bf16 matrix multiplication. A machine that fails compute is destroyed and replaced in 10 seconds. - Renter-Side Quarantine and Beta Reputation: Failed boots auto-quarantine the machine ID for a cooldown window so the retry loop cannot pick the same dud. Outcomes feed a per-machine time-decayed Beta posterior, reordering search results while preserving hard floors.
- Artifact Verification Before Destruction: The exported adapter or GGUF carries a SHA256 manifest. The local controller pulls, hashes the landed bytes, and asserts match before issuing the destroy command.
3. Asymmetric Stop and Destroy
Money is treated as a first-class failure mode. Automation is permitted to stop an instance (saving GPU costs while billing pennies for disk), but only the human operator is permitted to destroy an instance. A mistaken stop costs a few cents; a mistaken destroy loses a training run and its checkpoints.
Duplicate instance leaks are closed structurally: whenever an instance creation times out, the lease manager reconciles by label against the vendor API before any retry, preventing orphan instances from billing forever.
4. Comparison with SkyPilot and dstack
SkyPilot and dstack are excellent schedulers for multi-cloud placement across friendly, trustworthy providers. tuneharness solves the complementary problem: survival, fraud prevention, and honesty when renting cheap consumer GPUs from strangers.
| Capability | SkyPilot | dstack | tuneharness |
|---|---|---|---|
| Multi-cloud spot orchestration | Strong | Strong | None (by design) |
| Renter-side compute preflight (bf16 matmul) | No | No | Yes (10s fail-fast) |
| Frozen-pull watchdog (6 min kill) | Assumes datacenter | Assumes datacenter | Yes |
| Renter-side quarantine & Beta reputation | No | No | Yes |
| Ceiling a project's own config cannot raise | No | No | Yes (operator locked) |
| Append-only cost ledger surviving destroy | No | No | Yes |
| Artifact-hash-before-destroy guarantee | No | No | Yes |
| Deterministic deploy gate with canary tests | Out of scope | Out of scope | Yes |